Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.
Modalities
In / out price
$0.13 / $0.40
per 1M
Cached price
$0.05
per 1M
Context
256K
Max output
8K
Released
Apr 2, 2026
Providers
1 routeThe same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.
| Provider | Privacy | Input / 1M | Output / 1M | Cache read / 1M | Context | Max output | Uptime | Latency | Throughput |
|---|---|---|---|---|---|---|---|---|---|
| Private | $0.13 | $0.40 | $0.05 | 256K | 8K | — | — | — |
Privacy
2 / 4Private routing
Private tier
- Anonymous
- Private
- TEE
- E2EE
Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed. Learn more