Nvidia

NVIDIA Nemotron 3 Super 120B A12B

nvidia
Chat

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

Modalities

In / out price

$0.085 / $0.40

per 1M

Context

262K

Max output

16K

Released

Updated

Aug 2, 2026

Providers

1 route

The same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below.

ProviderPrivacyInput / 1MOutput / 1MCache read / 1MContextMax outputUptimeLatencyThroughput
DeepInfraDeepInfra
Private$0.085$0.40262K16K

Privacy

2 / 3

Private routing

Private tier

  1. Anonymous
  2. Private
  3. E2EE

Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed.

Features

3
Streaming
Tool calling
Reasoning

Uptime

NVIDIA Nemotron 3 Super 120B A12B — AnonRouter Models