Gemma

Google Gemma 4 31B Instruct

google
Chat

Gemma 4 31B is a dense model from Google DeepMind with 31B parameters, delivering frontier-level reasoning performance. It handles text, image, and video input, supports 256K context, function calling, and configurable thinking modes.

Modalities

In / out price

$0.13 / $0.38

per 1M

Cached price

$0.09

per 1M

Context

262K

Max output

16K

Released

Apr 3, 2026

Providers

3 routes

The same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.

ProviderPrivacyInput / 1MOutput / 1MCache read / 1MContextMax outputUptimeLatencyThroughput
TEE$0.40$1.00262K33K
VeniceVenice
Private$0.12$0.36$0.09256K8K
DeepInfraDeepInfra
Private$0.13$0.38262K16K

Privacy

3 / 4

Hardware-isolated routing

TEE tier

  1. Anonymous
  2. Private
  3. TEE
  4. E2EE

Inference runs inside a hardware-isolated, attestable environment. TEE protects processing, but does not by itself mean the client encrypted the request end to end. Learn more

Features

7
Streaming
Tool calling
Reasoning
Vision
JSON/schema
Web search
Prompt caching

Uptime

Google Gemma 4 31B Instruct — AnonRouter Models