Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.
Modalities
In / out price
$0.13 / $0.4
per 1M
Context
256K
Max output
8K
Released
Apr 2, 2026
Updated
Jul 17, 2026
Privacy
2 / 3Private routing
Private tier
- Anonymous
- Private
- E2EE
- Prompt and response content is not retained after the request.
- Request content is visible to the inference runtime while it is being processed.
Moderation
Not specified
No model-level moderation classification is recorded in this catalog.
Data retention
Zero retention
The provider reports that request content is discarded after inference.
Features
7Routing
2Anonrouter hosted
Routed through AnonRouter's gateway with metadata-only logging.
Provider direct
Requests egress directly to the provider runtime.
USD per 1M tokens from the Venice models API.