Nvidia

NVIDIA Nemotron 3 Ultra

nvidia
Chat

NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

Modalities

In / out price

$0.50 / $2.20

per 1M

Cached price

$0.10

per 1M

Context

262K

Max output

16K

Released

Jun 4, 2026

Providers

2 routes

The same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.

ProviderPrivacyInput / 1MOutput / 1MCache read / 1MContextMax outputUptimeLatencyThroughput
VeniceVenice
Private$0.625$3.125$0.1875256K33K
DeepInfraDeepInfra
Private$0.50$2.20$0.10262K16K

Privacy

2 / 4

Private routing

Private tier

  1. Anonymous
  2. Private
  3. TEE
  4. E2EE

Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed. Learn more

Features

6
Streaming
Tool calling
Reasoning
JSON/schema
Web search
Prompt caching

Uptime

NVIDIA Nemotron 3 Ultra — AnonRouter Models