Meta

Llama 4 Maverick 17B 128E Instruct FP8

meta-llama
Chat

The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Maverick, a 17 billion parameter model with 128 experts

Modalities

In / out price

$0.20 / $0.80

per 1M

Context

1M

Max output

16K

Released

Updated

Sep 6, 2026

Providers

1 route

The same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.

ProviderPrivacyInput / 1MOutput / 1MCache read / 1MContextMax outputUptimeLatencyThroughput
DeepInfraDeepInfra
Private$0.20$0.801M16K

Privacy

2 / 4

Private routing

Private tier

  1. Anonymous
  2. Private
  3. TEE
  4. E2EE

Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed. Learn more

Features

2
Streaming
Vision

Uptime

Llama 4 Maverick 17B 128E Instruct FP8 — AnonRouter Models