AI Models with Multi-Provider Failover

collections/multi-provider · 51 models

Models served by more than one provider. Routing policy picks between the routes per request, and a Private-only policy refuses to fall back to a weaker route rather than quietly downgrading the guarantee to keep the request alive.

Models

51

Labs

13

model creators

From (input)

$0.03

per 1M tokens

Max context

1M

Private routes

41 / 51

up to E2EE

In this collection51

Ordered by strongest privacy.

Zhipu
GLM 5.2
z-ai/glm-5.2
E2EE

GLM 5.2 running in a Trusted Execution Environment (TEE). Z.AI's flagship model for long-horizon tasks with enhanced reasoning and project-level engineering context, with hardware attestation evidence available for independent verification.

E2EE|524K context|$1.75/M input|$5.75/M output
Gemma
Gemma 4 26B A4B Uncensored
google/gemma-4-26b-a4b-uncensored
E2EE

Gemma 4 26B A4B Uncensored running in a Trusted Execution Environment (TEE). An uncensored variant of Google's Gemma 4 MoE model with 25.2B total / 3.8B active parameters, supporting multimodal input across text and images, with hardware attestation evidence available for independent verification.

E2EE|64K context|$0.19/M input|$0.88/M output
Qwen
Qwen 2.5 7B
qwen/qwen-2.5-7b
E2EE

Qwen 2.5 7B Instruct running in a Trusted Execution Environment (TEE). A compact model with strong coding, math, and multilingual capabilities supporting 29+ languages, with hardware attestation evidence available for independent verification.

E2EE|32K context|$0.05/M input|$0.13/M output
Gemma
Gemma 4 31B
google/gemma-4-31b-instruct
TEE

Gemma 4 31B served in a Tinfoil verified confidential enclave.

TEE|262K context|$0.40/M input|$1.00/M output
Kimi
Kimi K3
moonshotai/kimi-k3
TEE

Kimi K3 served in a Tinfoil verified confidential enclave.

TEE|262K context|$4.00/M input|$20.00/M output
Meta
Llama 3.3 70B
meta-llama/llama-3.3-70b
TEE

Llama 3.3 70B served in a Tinfoil verified confidential enclave.

TEE|131K context|$1.75/M input|$2.75/M output
OpenAI
OpenAI GPT OSS 120B
openai/gpt-oss-120b
TEE

OpenAI GPT OSS 120B served in a Tinfoil verified confidential enclave.

TEE|131K context|$0.15/M input|$0.60/M output
DeepSeek
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
Private

DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks.

Private|1M context|$0.09/M input|$0.18/M output
DeepSeek
DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
Private

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.

Private|1M context|$0.06/M input|$0.18/M output
Zhipu
GLM 5.3 Flash
z-ai/glm-5.3-flash
Private

GLM-5.3 Flash is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.

Private|1M context|$0.15/M input|$0.50/M output
DeepSeek
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
Private

DeepSeek V4 Pro is a 1.6T-parameter Mixture-of-Experts model with 49B active parameters and a 1M-token context window. Built for advanced reasoning, coding, and long-horizon agentic workflows with a hybrid attention system for efficient long-context processing.

Private|1M context|$1.65/M input|$3.301/M output
DeepSeek
DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813
Private

DeepSeek V4 Pro is a 1.6T-parameter Mixture-of-Experts model with 49B active parameters and a 1M-token context window. Built for advanced reasoning, coding, and long-horizon agentic workflows with a hybrid attention system for efficient long-context processing.

Private|1M context|$1.65/M input|$4.95/M output
DeepSeek
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash
Private

DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters (8B active on input, 16B on output) and a 1M-token context window. It natively processes images and text, with strong reasoning, coding, and agentic performance.

Private|1M context|$0.375/M input|$1.50/M output
Grok
Grok 4.3
x-ai/grok-4.3
Private

Grok 4.3 is xAI's most intelligent and fastest reasoning model with function calling, structured outputs, and a 1M-token context window. Suited for agentic workflows, instruction-following tasks, and applications requiring high factual accuracy.

Private|1M context|$1.42/M input|$2.83/M output
Venice
MiMo-V2.5
xiaomi/mimo-v2.5
Private

MiMo-V2.5 is Xiaomi's native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding in a unified architecture. Built on a sparse Mixture-of-Experts backbone with 310B total and 15B active parameters, it delivers long-context reasoning up to 1M tokens, function calling, and multimodal perception.

Private|1M context|$0.40/M input|$2.00/M output
Zhipu
GLM 5.3
z-ai/glm-5.3
Private

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

Private|1M context|$1.75/M input|$5.50/M output
Venice
Inkling
thinking-machines/inkling
Private

Inkling is a general-purpose multimodal model from Thinking Machines Lab that accepts text, image, and audio inputs and generates text. It is a 66-layer sparse MoE (975B total / 41B active) with hybrid local/global attention, 512K context, and variable thinking effort — suited for chat, coding, tool use, and agentic workflows. Video input is not supported on Venice.

Private|524K context|$1.25/M input|$5.0625/M output
Qwen
Qwen3.5 397B A17B
qwen/qwen-3.5-397b
Private

Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks.

Private|262K context|$0.45/M input|$3.00/M output
Qwen
Qwen 3.8 2.4T
qwen/qwen-3.8-2.4t
Private

Qwen 3.8 2.4T is Alibaba's open-weight 2.4-trillion-parameter MoE model (95B active), with major gains in software engineering, research, and long-horizon agentic tasks. It is text-only, requires thinking mode, and supports a 262K-token context window.

Private|262K context|$2.50/M input|$7.50/M output
Qwen
Qwen 3.8 27B
qwen/qwen-3.8-27b
Private

Qwen 3.8 27B is a native vision-language dense model with 27B parameters. It improves coding, professional work, research, and long-horizon agentic tasks, with flexible thinking control and image and video understanding. It supports a native 262K-token context window.

Private|262K context|$0.45/M input|$3.20/M output
Qwen
Qwen3 VL 30B A3B Instruct
qwen/qwen3-vl-30b-a3b
Private

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities.

Private|262K context|$0.15/M input|$0.60/M output
Kimi
Kimi K2.6
moonshotai/kimi-k2.6
Private

Kimi K2.6 is an open-source, native multimodal agentic model from Moonshot AI with 1T total parameters and 32B active parameters. It excels at long-horizon coding, coding-driven design, agent swarm orchestration, and proactive autonomous execution with 256K context windows.

Private|256K context|$0.75/M input|$3.50/M output
Kimi
Kimi K2.7 Code
moonshotai/kimi-k2.7-code
Private

Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built on Kimi K2.6, with 1T total parameters and 32B active parameters. It always operates in thinking mode, supports text and image input, and targets long-horizon software engineering, agentic task decomposition, and multi-turn coding workflows with 256K context.

Private|256K context|$0.75/M input|$3.50/M output
Nvidia
NVIDIA Nemotron 3 Ultra
nvidia/nvidia-nemotron-3-ultra
Private

NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

Private|256K context|$0.625/M input|$3.125/M output
Qwen
Qwen 3 Coder 480B Turbo
qwen/qwen-3-coder-480b-turbo
Private

Turbo variant of Qwen3 Coder 480B, optimized for faster inference on code tasks.

Private|256K context|$0.35/M input|$1.50/M output
Qwen
Qwen 3 Next 80b
qwen/qwen-3-next-80b
Private

Optimized for speed and efficiency.

Private|256K context|$0.35/M input|$1.90/M output
Qwen
Qwen 3.5 35B A3B
qwen/qwen-3.5-35b-a3b
Private

Qwen 3.5 35B A3B is a highly efficient MoE model with 35B total parameters and only 3B active parameters. It surpasses the larger Qwen3-235B-A22B while being 6.7x smaller, excelling at reasoning, coding, and general knowledge tasks.

Private|256K context|$0.3125/M input|$1.25/M output
Qwen
Qwen 3.5 9B
qwen/qwen-3.5-9b
Private

A 9B dense model with 262K native context window (extendable to 1M). Features Gated DeltaNet hybrid attention architecture for efficient long-context processing. Supports 201 languages, thinking/reasoning mode, and function calling.

Private|256K context|$0.10/M input|$0.15/M output
Qwen
Qwen 3.6 27B
qwen/qwen-3.6-27b
Private

The Qwen 3.6 27B native vision-language dense model builds upon the 3.5-27B version, with key improvements in agentic coding capabilities and enhanced STEM reasoning and inference skills. In the vision modality, it demonstrates significant advances in spatial intelligence, object localization, and detection, while video understanding, document OCR, and visual agent capabilities continue to improve steadily.

Private|256K context|$0.325/M input|$3.25/M output
Qwen
Qwen 3.6 35B A3B
qwen/qwen-3.6-35b-a3b
Private

Qwen 3.6 35B A3B is a fast mixture-of-experts model with 35B total parameters and ~3B active per token. Strong at agentic coding, STEM reasoning, and tool use, with a native 256K context window.

Private|256K context|$0.10/M input|$1.00/M output
Zhipu
GLM 5.1
z-ai/glm-5.1
Private

GLM-5.1 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple languages. Supports large context windows for processing extensive text and detailed analysis with fast inference speed.

Private|200K context|$1.54/M input|$4.84/M output
Zhipu
GLM 4.6
z-ai/glm-4.6
Private

GLM-4.6 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.

Private|198K context|$0.43/M input|$1.75/M output
Zhipu
GLM 4.7
z-ai/glm-4.7
Private

GLM-4.7 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.

Private|198K context|$0.55/M input|$2.65/M output
DeepSeek
DeepSeek V3.2
deepseek/deepseek-v3.2
Private

DeepSeek-V3.2 is an efficient large language model with DeepSeek Sparse Attention (DSA) for long contexts. It features strong reasoning and tool-use skills, achieving top results on the 2025 IMO and IOI.

Private|160K context|$0.33/M input|$0.48/M output
Gemma
gemma 3 4b it
google/gemma-3-4b-it
Private

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3-12B is Google's latest open source model, successor to Gemma 2

Private|131K context|$0.05/M input|$0.10/M output
Meta
Muse Glimmer 30B
meta-models/muse-glimmer-30b
Private

Muse Glimmer is a 30B multimodal agentic model distilled from Muse Spark — reasoning, tool use, and failure recovery in a single model that runs locally on consumer hardware.

Private|131K context|$0.30/M input|$1.20/M output
OpenAI
OpenAI GPT OSS 20B
openai/gpt-oss-20b
Private

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.

Private|131K context|$0.03/M input|$0.14/M output
Meta
Hermes 3 Llama 3.1 405b
meta-llama/hermes-3-llama-3.1-405b
Private

Hermes 3 405B is a frontier level, full parameter finetune of the Llama-3.1 405B foundation model, focused on aligning LLMs to the user, with powerful steering capabilities and control given to the end user.

Private|128K context|$1.10/M input|$3.00/M output
Qwen
Qwen 3 235B A22B Instruct 2507
qwen/qwen-3-235b-a22b-instruct-2507
Private

Built for in-depth research and handling long, complex documents. Ideal for technical work, multimodal input, and high-precision tasks.

Private|128K context|$0.15/M input|$0.75/M output
Qwen
Qwen3 VL 235B
qwen/qwen3-vl-235b
Private

Qwen3-VL 235B vision-language model with MoE architecture. The most powerful VL model in the Qwen series with superior visual perception, OCR, and multimodal reasoning.

Private|128K context|$0.21/M input|$1.90/M output
Qwen
Qwen3 32B
qwen/qwen3-32b
Private

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support

Private|41K context|$0.08/M input|$0.28/M output
Claude
Claude Fable 5
anthropic/claude-fable-5
Anonymous

Claude Fable 5 is Anthropic's most capable widely released model, designed for demanding reasoning and long-horizon agentic work. It features a 1M token context window, 128K max output tokens, always-on adaptive thinking, and strong multimodal capabilities.

Anonymous|1M context|$12.00/M input|$60.00/M output
Claude
Claude Opus 4.7
anthropic/claude-opus-4.7
Anonymous

Claude Opus 4.7 is Anthropic's most capable generally available model for complex reasoning and agentic coding. It features a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.

Anonymous|1M context|$6.00/M input|$30.00/M output
Claude
Claude Opus 4.8
anthropic/claude-opus-4.8
Anonymous

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports long-horizon agentic work, complex multi-step coding, and memory-driven tasks where coherence over extended sessions matters. It features a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.

Anonymous|1M context|$6.00/M input|$30.00/M output
Claude
Claude Opus 5
anthropic/claude-opus-5
Anonymous

Claude Opus 5 is Anthropic's most capable model in the Opus family. It delivers major gains over Opus 4.8 in agentic coding, professional knowledge work, and long-horizon reasoning, with a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.

Anonymous|1M context|$6.00/M input|$30.00/M output
Claude
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anonymous

Claude Sonnet 4.6 is Anthropic's best combination of speed and intelligence, offering strong performance on coding, reasoning, and general tasks with excellent speed and cost efficiency. It features a 1M token context window and 64K max output tokens.

Anonymous|1M context|$3.60/M input|$18.00/M output
Claude
Claude Sonnet 5
anthropic/claude-sonnet-5
Anonymous

Claude Sonnet 5 is Anthropic's latest Sonnet model, substantially improving on Sonnet 4.6 in coding and agentic work and reaching near-Opus quality on many tasks. It features a 1M token context window, adaptive thinking, and strong document and vision understanding.

Anonymous|1M context|$3.00/M input|$15.00/M output
Gemini
Gemini 3.5 Flash
google/gemini-3.5-flash
Anonymous

Gemini 3.5 Flash is a high speed, high value thinking model with 1M context, designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro level reasoning with substantially lower latency.

Anonymous|1M context|$1.55/M input|$9.45/M output
Gemini
Gemini 3.7 Flash
google/gemini-3.7-flash
Anonymous

Gemini 3.7 Flash is Google's most capable Flash model, built for complex coding, agentic workflows, and reliable multi-step execution, with 1M context and tunable thinking.

Anonymous|1M context|$0.9375/M input|$4.6875/M output
Qwen
Qwen 3.7 Max
qwen/qwen-3.7-max
Anonymous

Qwen 3.7 Max is the largest model in the Qwen 3.7 series, with deep thinking, function calling, prompt caching, and multimodal input support for images and video. It excels at programming, office and productivity tasks, and long-running autonomous agent workflows.

Anonymous|1M context|$2.70/M input|$8.05/M output
Qwen
Qwen 3.8 Max
qwen/qwen-3.8-max
Anonymous

Qwen 3.8 Max is Alibaba's flagship 2.4-trillion-parameter MoE model, with major gains over Qwen 3.7 Max in software engineering and office-productivity workflows and strong long-horizon, multi-agent performance. It accepts both text and vision-language input (images and video), operates in thinking mode only, and supports a 1M-token context window.

Anonymous|1M context|$2.50/M input|$7.50/M output

This list is rebuilt from the live catalog rather than stored as a snapshot, so it tracks pricing, context windows, and privacy tiers as providers change them. Ordering is yours to pick, and there is no popularity option: prompts are never retained, and the usage metadata kept for billing is not turned into a public ranking.

Explore more collections

Strongest guarantee in this collection:E2EE

AI Models with Multi-Provider Failover - AnonRouter