NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.
Modalities
In / out price
$0.085 / $0.40
per 1M
Context
262K
Max output
16K
Released
–
Updated
Aug 2, 2026
Providers
1 routeThe same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below.
| Provider | Privacy | Input / 1M | Output / 1M | Cache read / 1M | Context | Max output | Uptime | Latency | Throughput |
|---|---|---|---|---|---|---|---|---|---|
| Private | $0.085 | $0.40 | — | 262K | 16K | — | — | — |
Privacy
2 / 3Private routing
Private tier
- Anonymous
- Private
- E2EE
Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed.