Nemotron Cascade 2 30B A3B is a reasoning-optimized language model from NVIDIA, designed for efficient inference with strong reasoning capabilities across complex tasks.
Beta
Modalities
In / out price
$0.14 / $0.8
per 1M
Context
256K
Max output
33K
Released
Mar 24, 2026
Updated
Jul 17, 2026
Privacy
2 / 3Private routing
Private tier
- Anonymous
- Private
- E2EE
- Prompt and response content is not retained after the request.
- Request content is visible to the inference runtime while it is being processed.
Moderation
Not specified
No model-level moderation classification is recorded in this catalog.
Data retention
Zero retention
The provider reports that request content is discarded after inference.
Features
5Streaming
Tool calling
Reasoning
JSON/schema
Web search
Routing
2Anonrouter hosted
Routed through AnonRouter's gateway with metadata-only logging.
Provider direct
Requests egress directly to the provider runtime.
USD per 1M tokens from the Venice models API.