DeepInfra

DeepInfra

inference provider · 13 models

Access 13 models served through DeepInfra on AnonRouter's privacy-first gateway, including DeepSeek V3.2, DeepSeek V4 Flash, and DeepSeek V4 Pro. DeepInfra says open-model inputs and outputs stay in memory only for the request and are deleted afterward, logging metadata rather than content; AnonRouter uses its standard Chat Completions path and excludes DeepInfra's retaining partner routes.

Models

13

Modalities

1

Text

From (input)

$0.03

per 1M tokens

Max context

1M

Private routes

13 / 13

not anonymous-only

Catalog by modality

13 routes
Text13

DeepInfra models13

DeepSeek
DeepSeek V3.2
deepseek/deepseek-v3.2
Private

DeepSeek-V3.2 is an efficient large language model with DeepSeek Sparse Attention (DSA) for long contexts. It features strong reasoning and tool-use skills, achieving top results on the 2025 IMO and IOI.

Private|164K context|$0.26/M input|$0.38/M output
DeepSeek
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
Private

DeepSeek V4 Flash running in a Trusted Execution Environment (TEE). Hardware attestation evidence is available for independent verification of enclave identity and configuration.

Private|1M context|$0.09/M input|$0.18/M output
DeepSeek
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
Private

DeepSeek V4 Pro is a 1.6T-parameter Mixture-of-Experts model with 49B active parameters and a 1M-token context window. Built for advanced reasoning, coding, and long-horizon agentic workflows with a hybrid attention system for efficient long-context processing.

Private|1M context|$1.30/M input|$2.60/M output
Gemma
Google Gemma 4 31B Instruct
google/google-gemma-4-31b-instruct
Private

Gemma 4 31B is a dense model from Google DeepMind with 31B parameters, delivering frontier-level reasoning performance. It handles text, image, and video input, supports 256K context, function calling, and configurable thinking modes.

Private|262K context|$0.13/M input|$0.38/M output
Kimi
Kimi K2.5
moonshotai/kimi-k2.5
Private

Kimi K2.5 is Moonshot AIs most advanced open reasoning model, featuring trillion-parameter Mixture-of-Experts architecture with 32B active parameters and 256K context windows.

Private|262K context|$0.45/M input|$2.25/M output
Kimi
Kimi K2.6
moonshotai/kimi-k2.6
Private

Kimi K2.6 is an open-source, native multimodal agentic model from Moonshot AI with 1T total parameters and 32B active parameters. It excels at long-horizon coding, coding-driven design, agent swarm orchestration, and proactive autonomous execution with 256K context windows.

Private|262K context|$0.75/M input|$3.50/M output
Kimi
Kimi K2.7 Code
moonshotai/kimi-k2.7-code
Private

Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built on Kimi K2.6, with 1T total parameters and 32B active parameters. It always operates in thinking mode, supports text and image input, and targets long-horizon software engineering, agentic task decomposition, and multi-turn coding workflows with 256K context.

Private|262K context|$0.74/M input|$3.50/M output
Nvidia
NVIDIA Nemotron 3 Ultra
nvidia/nvidia-nemotron-3-ultra
Private

NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

Private|262K context|$0.50/M input|$2.20/M output
OpenAI
GPT OSS 20B
openai/gpt-oss-20b
Private

GPT OSS 20B running in a Trusted Execution Environment (TEE). OpenAI's compact open-weight 21B MoE model with 3.6B active parameters, optimized for lower-latency inference, with hardware attestation evidence available for independent verification.

Private|131K context|$0.03/M input|$0.14/M output
Zhipu
GLM 5.1
z-ai/glm-5.1
Private

GLM 5.1 running in a Trusted Execution Environment (TEE). Hardware attestation evidence is available for independent verification of enclave identity and configuration.

Private|203K context|$1.05/M input|$3.50/M output
Zhipu
GLM 5.2
z-ai/glm-5.2
Private

GLM 5.2 running in a Trusted Execution Environment (TEE). Z.AI's flagship model for long-horizon tasks with enhanced reasoning and project-level engineering context, with hardware attestation evidence available for independent verification.

Private|1M context|$0.75/M input|$2.40/M output
Nvidia
NVIDIA Nemotron 3 Super 120B A12B
nvidia/nvidia-nemotron-3-super-120b-a12b
Private

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

Private|262K context|$0.085/M input|$0.40/M output
Qwen
Qwen3.6 35B A3B
qwen/qwen3.6-35b-a3b
Private

Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.

Private|262K context|$0.10/M input|$0.95/M output
Provider documentation3 sources
DeepInfra provider — AnonRouter