Gemma

Google Gemma 4 26B A4B Instruct

google
Chat

Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.

Modalities

In / out price

$0.13 / $0.40

per 1M

Cached price

$0.05

per 1M

Context

256K

Max output

8K

Released

Apr 2, 2026

Providers

1 route

The same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.

ProviderPrivacyInput / 1MOutput / 1MCache read / 1MContextMax outputUptimeLatencyThroughput
VeniceVenice
Private$0.13$0.40$0.05256K8K

Privacy

2 / 4

Private routing

Private tier

  1. Anonymous
  2. Private
  3. TEE
  4. E2EE

Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed. Learn more

Features

7
Streaming
Tool calling
Reasoning
Vision
JSON/schema
Web search
Prompt caching

Uptime

Google Gemma 4 26B A4B Instruct — AnonRouter Models