Llama 3.3 70B Instruct Turbo
Llama 3.3-70B Turbo is a highly optimized version of the Llama 3.3-70B model, utilizing FP8 quantization to deliver significantly faster inference speeds with a minor trade-off in accuracy. The model is designed to be helpful, safe, and flexible, with a focus on responsible deployment and mitigating potential risks such as bias, toxicity, and misinformation. It achieves state-of-the-art performance on various benchmarks, including conversational tasks, language translation, and text generation.
Modalities
In / out price
$0.10 / $0.32
per 1M
Context
131K
Max output
16K
Released
–
Updated
Sep 6, 2026
Providers
1 routeThe same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.
| Provider | Privacy | Input / 1M | Output / 1M | Cache read / 1M | Context | Max output | Uptime | Latency | Throughput |
|---|---|---|---|---|---|---|---|---|---|
| Private | $0.10 | $0.32 | — | 131K | 16K | — | — | — |
Privacy
2 / 4Private routing
Private tier
- Anonymous
- Private
- TEE
- E2EE
Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed. Learn more