Nemotron Cascade 2 30B A3B is a reasoning-optimized language model from NVIDIA, designed for efficient inference with strong reasoning capabilities across complex tasks.
NVIDIA
nvidia/ · 9 models
Access 9 NVIDIA models through AnonRouter's privacy-first gateway, including Nemotron Cascade 2 30B A3B, NVIDIA Nemotron 3 Nano 30B, and NVIDIA Nemotron 3 Ultra. Compare pricing, context windows, and capabilities across NVIDIA's text, audio, embeddings routes — every request anonymized, with no payload logging.
Models
9
Modalities
3
Text, Audio, Embeddings
From (input)
$0.0125
per 1M tokens
Max context
262K
Private routes
9 / 9
not anonymous-only
Catalog by modality
9 routesNVIDIA models9
NVIDIA Nemotron 3 Nano 30B is a compact and efficient language model from NVIDIA, optimized for fast inference while maintaining strong performance across diverse tasks.
NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.
Nemotron Nano 9B V2 served through Amazon Bedrock with support for AWS zero-data-retention mode.
Nemotron Nano 12B V2 served through Amazon Bedrock with support for AWS zero-data-retention mode.
Nemotron Super 3 120B served through Amazon Bedrock with support for AWS zero-data-retention mode.
NVIDIA Nemotron 3 Super 120B A12B served on DeepInfra serverless inference.
Embedding model for semantic search and retrieval, served through Venice.
Speech-to-text transcription model served through Venice.