NVIDIA Nemotron 3 Nano 30B is a compact and efficient language model from NVIDIA, optimized for fast inference while maintaining strong performance across diverse tasks.
Beta
Modalities
In / out price
$0.075 / $0.3
per 1M
Context
128K
Max output
16K
Released
Jan 27, 2026
Updated
Jul 17, 2026
Privacy
2 / 3Private routing
Private tier
- Anonymous
- Private
- E2EE
- Prompt and response content is not retained after the request.
- Request content is visible to the inference runtime while it is being processed.
Moderation
Not specified
No model-level moderation classification is recorded in this catalog.
Data retention
Zero retention
The provider reports that request content is discarded after inference.
Features
4Streaming
Tool calling
JSON/schema
Web search
Routing
2Anonrouter hosted
Routed through AnonRouter's gateway with metadata-only logging.
Provider direct
Requests egress directly to the provider runtime.
USD per 1M tokens from the Venice models API.