Venice

Venice

inference provider · 288 models

Access 288 models served through Venice on AnonRouter's privacy-first gateway, including Aion 3.0, Aion 3.0 Mini, and Claude Fable 5. Venice serves privacy-first inference with no payload logging, and AnonRouter strips identity before requests ever reach it.

Models

288

Modalities

5

Text, Video, Image, Audio, Embeddings

From (input)

$0.0125

per 1M tokens

Max context

2M

Private routes

106 / 288

not anonymous-only

Catalog by modality

288 routes
Text101Video96Image54Audio28Embeddings9

Venice models288

Aion 3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. Multiple specialized models collaborate on each response to produce stronger narrative…

128K ctx·$3.75/M in·$7.5/M out·Updated Jul 17, 2026Anonymous

Aion 3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. Multiple specialized models collaborate on each response to produce stronger…

128K ctx·$0.875/M in·$1.75/M out·Updated Jul 17, 2026Anonymous

Claude Fable 5 is Anthropic's most capable widely released model, designed for demanding reasoning and long-horizon agentic work. It features a 1M token context window, 128K max output tokens,…

1M ctx·$12/M in·$60/M out·Updated Jul 17, 2026Anonymous

Claude Opus 4.5 is Anthropic's frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities,…

198K ctx·$6/M in·$30/M out·Updated Jul 17, 2026Anonymous

Claude Opus 4.6 is Anthropic's most capable reasoning model, building on Opus 4.5 with enhanced performance across complex software engineering, agentic workflows, and long-horizon tasks. It features…

1M ctx·$6/M in·$30/M out·Updated Jul 17, 2026Anonymous

Claude Opus 4.7 is Anthropic's most capable generally available model for complex reasoning and agentic coding. It features a 1M token context window, 128K max output tokens, adaptive thinking, and…

1M ctx·$6/M in·$30/M out·Updated Jul 17, 2026Anonymous

Claude Opus 4.7 (Fast) is a speed-optimized variant of Anthropic's most capable generally available model, offering the same 1M token context window and strong performance across complex reasoning…

1M ctx·$36/M in·$180/M out·Updated Jul 17, 2026Anonymous

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports long-horizon agentic work, complex multi-step coding, and memory-driven tasks where coherence…

1M ctx·$6/M in·$30/M out·Updated Jul 17, 2026Anonymous

Claude Opus 4.8 (Fast) is a speed-optimized variant of Anthropic's most capable generally available Opus model, offering the same 1M token context window and strong performance across long-horizon…

1M ctx·$12/M in·$60/M out·Updated Jul 17, 2026Anonymous

Claude Sonnet 4.5 is Anthropic's balanced model offering strong performance on coding, reasoning, and general tasks with good speed and cost efficiency.

198K ctx·$3.75/M in·$18.75/M out·Updated Jul 17, 2026Anonymous

Claude Sonnet 4.6 is Anthropic's best combination of speed and intelligence, offering strong performance on coding, reasoning, and general tasks with excellent speed and cost efficiency. It features…

1M ctx·$3.6/M in·$18/M out·Updated Jul 17, 2026Anonymous

Claude Sonnet 5 is Anthropic's latest Sonnet model, substantially improving on Sonnet 4.6 in coding and agentic work and reaching near-Opus quality on many tasks. It features a 1M token context…

1M ctx·$3/M in·$15/M out·Updated Jul 17, 2026Anonymous

DeepSeek-V3.2 is an efficient large language model with DeepSeek Sparse Attention (DSA) for long contexts. It features strong reasoning and tool-use skills, achieving top results on the 2025 IMO and…

160K ctx·$0.33/M in·$0.48/M out·Updated Jul 17, 2026Private

DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads…

1M ctx·$0.138/M in·$0.275/M out·Updated Jul 17, 2026Anonymous

DeepSeek V4 Pro is a 1.6T-parameter Mixture-of-Experts model with 49B active parameters and a 1M-token context window. Built for advanced reasoning, coding, and long-horizon agentic workflows with a…

1M ctx·$1.65/M in·$3.301/M out·Updated Jul 17, 2026Anonymous

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro level reasoning with substantially lower…

256K ctx·$0.7/M in·$3.75/M out·Updated Jul 17, 2026Anonymous

Gemini 3.5 Flash is a high speed, high value thinking model with 1M context, designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro level reasoning with…

1M ctx·$1.55/M in·$9.45/M out·Updated Jul 17, 2026Anonymous

Gemma 3 27B running in a Trusted Execution Environment (TEE). Google's multimodal model supporting vision-language input with 140+ language understanding, with hardware attestation evidence available…

40K ctx·$0.14/M in·$0.5/M out·Updated Jul 17, 2026E2EE

Gemma 4 Uncensored is an uncensored variant of Google Gemma 4 26B, a Mixture-of-Experts model with 26B total parameters and only 4B active per token. Fine-tuned for uncensored chat without content…

256K ctx·$0.1625/M in·$0.5/M out·Updated Jul 17, 2026Private

Mercury 2 is a diffusion-based reasoning LLM from Inception, delivering over 1,000 tokens per second — 5x faster than leading speed-optimized models — with strong reasoning, tool use, and structured…

128K ctx·$0.3125/M in·$0.9375/M out·Updated Jul 17, 2026Anonymous

Inkling is a general-purpose multimodal model from Thinking Machines Lab that accepts text, image, and audio inputs and generates text. It is a 66-layer sparse MoE (975B total / 41B active) with…

1M ctx·$2.3375/M in·$5.85/M out·Updated Jul 17, 2026Private

MiniMax-M2.5 is a state-of-the-art large language model optimized for coding, agentic workflows, and modern application development with enhanced reasoning capabilities.

198K ctx·$0.27/M in·$0.95/M out·Updated Jul 17, 2026Private

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity with advanced agentic capabilities through multi-agent collaboration.

198K ctx·$0.375/M in·$1.5/M out·Updated Jul 17, 2026Private

MiniMax-M3 preview is a 1.4T-parameter frontier model from MiniMax for coding, agentic workflows, and complex reasoning, served at fp8 with a 512K context window.

524K ctx·$0.3/M in·$1.2/M out·Updated Jul 17, 2026Private

Mistral Small 4 unifies instruction following, reasoning, coding, and vision in a single 119B MoE model with 256K context and configurable reasoning effort.

256K ctx·$0.1875/M in·$0.75/M out·Updated Jul 17, 2026Private

Kimi K2.5 is Moonshot AIs most advanced open reasoning model, featuring trillion-parameter Mixture-of-Experts architecture with 32B active parameters and 256K context windows.

256K ctx·$0.56/M in·$3.5/M out·Updated Jul 17, 2026Private

Kimi K2.6 is an open-source, native multimodal agentic model from Moonshot AI with 1T total parameters and 32B active parameters. It excels at long-horizon coding, coding-driven design, agent swarm…

256K ctx·$0.75/M in·$3.5/M out·Updated Jul 17, 2026Private

Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built on Kimi K2.6, with 1T total parameters and 32B active parameters. It always operates in thinking mode, supports text and image…

256K ctx·$0.75/M in·$3.5/M out·Updated Jul 17, 2026Private

Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly…

1M ctx·$3.75/M in·$18.75/M out·Updated Jul 17, 2026Anonymous

NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost…

256K ctx·$0.625/M in·$3.125/M out·Updated Jul 17, 2026Private

OpenAI's multimodal flagship model with vision capabilities, strong reasoning, and broad knowledge. Popular for its balanced performance across tasks. Version: 2024-11-20.

128K ctx·$3.125/M in·$12.5/M out·Updated Jul 17, 2026Anonymous

OpenAI's cost-efficient small model that delivers GPT-4 level intelligence at a fraction of the cost. Ideal for high-volume applications requiring strong reasoning. Version: 2024-07-18.

128K ctx·$0.1875/M in·$0.75/M out·Updated Jul 17, 2026Anonymous

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context performance compared to GPT-5.1. It uses adaptive reasoning to allocate computation…

256K ctx·$2.19/M in·$17.5/M out·Updated Jul 17, 2026Anonymous

GPT-5.2 Codex is OpenAI specialized coding model built on GPT-5.2, optimized for advanced software development, code generation, and technical problem-solving.

256K ctx·$2.19/M in·$17.5/M out·Updated Jul 17, 2026Anonymous

GPT-5.3 Codex is OpenAI specialized coding model built on GPT-5.3, optimized for advanced software development, code generation, and technical problem-solving.

400K ctx·$2.19/M in·$17.5/M out·Updated Jul 17, 2026Anonymous

GPT-5.4 is the latest frontier model in the GPT-5 series with a 1M+ context window, offering improved agentic and long context performance. It uses adaptive reasoning to dynamically allocate…

1M ctx·$3.13/M in·$18.8/M out·Updated Jul 17, 2026Anonymous

GPT-5.4 Mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across…

400K ctx·$0.9375/M in·$5.625/M out·Updated Jul 17, 2026Anonymous

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning for complex, high-stakes tasks. It provides a 1M+ token context window (922K input,…

1M ctx·$37.5/M in·$225/M out·Updated Jul 17, 2026Anonymous

GPT-5.5 is the latest frontier model in the GPT-5 series with a 1M+ context window, offering improved agentic and long context performance. It uses adaptive reasoning to dynamically allocate…

1M ctx·$6.25/M in·$37.5/M out·Updated Jul 17, 2026Anonymous

GPT-5.5 Pro is OpenAI's most advanced model, building on GPT-5.5's unified architecture with enhanced reasoning for complex, high-stakes tasks. It provides a 1M+ token context window (922K input,…

1M ctx·$37.5/M in·$225/M out·Updated Jul 17, 2026Anonymous

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows,…

1M ctx·$1.25/M in·$7.5/M out·Updated Jul 17, 2026Anonymous

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

1M ctx·$1.25/M in·$7.5/M out·Updated Jul 17, 2026Anonymous

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks…

1M ctx·$6.25/M in·$37.5/M out·Updated Jul 17, 2026Anonymous

GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

1M ctx·$6.25/M in·$37.5/M out·Updated Jul 17, 2026Anonymous

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks…

1M ctx·$3.125/M in·$18.75/M out·Updated Jul 17, 2026Anonymous

GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

1M ctx·$3.125/M in·$18.75/M out·Updated Jul 17, 2026Anonymous

GPT OSS 120B running in a Trusted Execution Environment (TEE). OpenAI's open-weight 117B-parameter MoE model with configurable reasoning depth and native tool use, with hardware attestation evidence…

128K ctx·$0.13/M in·$0.65/M out·Updated Jul 17, 2026E2EE

GPT OSS 20B running in a Trusted Execution Environment (TEE). OpenAI's compact open-weight 21B MoE model with 3.6B active parameters, optimized for lower-latency inference, with hardware attestation…

128K ctx·$0.05/M in·$0.19/M out·Updated Jul 17, 2026E2EE

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. The model supports…

128K ctx·$0.07/M in·$0.3/M out·Updated Jul 17, 2026Private

Qwen 2.5 7B Instruct running in a Trusted Execution Environment (TEE). A compact model with strong coding, math, and multilingual capabilities supporting 29+ languages, with hardware attestation…

32K ctx·$0.05/M in·$0.13/M out·Updated Jul 17, 2026E2EE

Qwen 3.5 35B A3B is a highly efficient MoE model with 35B total parameters and only 3B active parameters. It surpasses the larger Qwen3-235B-A22B while being 6.7x smaller, excelling at reasoning,…

256K ctx·$0.3125/M in·$1.25/M out·Updated Jul 17, 2026Private

Qwen 3.5 is Alibaba flagship reasoning model featuring a 397B parameter Mixture-of-Experts architecture with 17B active parameters. It excels at complex reasoning, coding, and general knowledge tasks.

128K ctx·$0.75/M in·$4.5/M out·Updated Jul 17, 2026Anonymous

A 9B dense model with 262K native context window (extendable to 1M). Features Gated DeltaNet hybrid attention architecture for efficient long-context processing. Supports 201 languages,…

256K ctx·$0.1/M in·$0.15/M out·Updated Jul 17, 2026Private

The Qwen 3.6 27B native vision-language dense model builds upon the 3.5-27B version, with key improvements in agentic coding capabilities and enhanced STEM reasoning and inference skills. In the…

256K ctx·$0.325/M in·$3.25/M out·Updated Jul 17, 2026Private

Qwen 3.6 27B FP8 running in a Trusted Execution Environment (TEE). Hardware attestation evidence is available for independent verification of enclave identity and configuration.

256K ctx·$0.346/M in·$3.46/M out·Updated Jul 17, 2026E2EE

Qwen 3.6 35B A3B FP8 running in a Trusted Execution Environment (TEE). A fast mixture-of-experts model with ~3B active parameters per token. Hardware attestation evidence is available for independent…

32K ctx·$0.182/M in·$1.18/M out·Updated Jul 17, 2026E2EE

Qwen 3.6 Plus Uncensored is Alibaba's latest flagship reasoning model with exceptional performance across coding, reasoning, and general knowledge tasks. Features mixed reasoning, function calling,…

1M ctx·$0.625/M in·$3.75/M out·Updated Jul 17, 2026Anonymous

Qwen 3.7 Max is the largest model in the Qwen 3.7 series, with deep thinking, function calling, prompt caching, and multimodal input support for images and video. It excels at programming, office and…

1M ctx·$2.7/M in·$8.05/M out·Updated Jul 17, 2026Anonymous

Qwen 3.7 Plus is Alibaba's latest flagship reasoning model with exceptional performance across coding, reasoning, and general knowledge tasks. Features mixed reasoning, function calling, and…

1M ctx·$0.5/M in·$2/M out·Updated Jul 17, 2026Anonymous

Qwen3 30B A3B running in a Trusted Execution Environment (TEE). A MoE model with 30.5B total parameters and 3.3B activated per inference, supporting ultra-long 256K context, with hardware attestation…

256K ctx·$0.19/M in·$0.69/M out·Updated Jul 17, 2026E2EE

Qwen3-VL 235B vision-language model with MoE architecture. The most powerful VL model in the Qwen series with superior visual perception, OCR, and multimodal reasoning.

128K ctx·$0.21/M in·$1.9/M out·Updated Jul 17, 2026Private

Qwen3 VL 30B A3B running in a Trusted Execution Environment (TEE). A multimodal model unifying text generation with visual understanding for images and videos, with hardware attestation evidence…

128K ctx·$0.25/M in·$0.9/M out·Updated Jul 17, 2026E2EE

Venice Uncensored 1.2 is designed for maximum creative freedom and authentic interaction. Built for open-ended exploration, roleplay, and unfiltered dialogue with improved capabilities over 1.1.

128K ctx·$0.2/M in·$0.9/M out·Updated Jul 17, 2026Private

Grok 4.20 is xAI's latest multimodal reasoning model with strong tool use, structured output support, and a 2M-token context window.

2M ctx·$1.42/M in·$2.83/M out·Updated Jul 17, 2026Private

Grok 4.20 Multi-Agent is a variant of xAI Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and…

2M ctx·$1.42/M in·$2.83/M out·Updated Jul 17, 2026Private

Grok 4.3 is xAI's most intelligent and fastest reasoning model with function calling, structured outputs, and a 1M-token context window. Suited for agentic workflows, instruction-following tasks, and…

1M ctx·$1.42/M in·$2.83/M out·Updated Jul 17, 2026Private

Grok 4.5 is xAI's intelligent coding model for agentic software engineering and workflow tasks, with function calling, structured outputs, and a 500K-token context window.

500K ctx·$2.27/M in·$6.8/M out·Updated Jul 17, 2026Private

MiMo-V2.5 is Xiaomi's native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding in a unified architecture. Built on a sparse Mixture-of-Experts…

1M ctx·$0.14/M in·$0.28/M out·Updated Jul 17, 2026Private

GLM-4.6 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive…

198K ctx·$0.43/M in·$1.75/M out·Updated Jul 17, 2026Private

GLM-4.7 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive…

198K ctx·$0.55/M in·$2.65/M out·Updated Jul 17, 2026Private

GLM 4.7 running in a Trusted Execution Environment (TEE). Z.AI's flagship model with enhanced programming capabilities and stable multi-step reasoning, with hardware attestation evidence available…

128K ctx·$1.1/M in·$4.15/M out·Updated Jul 17, 2026E2EE

GLM-4.7-Flash is a fast inference variant of GLM-4.7, optimized for speed while maintaining strong reasoning capabilities. Ideal for applications requiring quick responses with good quality.

128K ctx·$0.125/M in·$0.5/M out·Updated Jul 17, 2026Private

GLM-5 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple languages.…

198K ctx·$1/M in·$3.2/M out·Updated Jul 17, 2026Private

GLM-5 Turbo is a fast inference model from Z.ai tuned for strong performance in agent-driven environments and production coding workflows.

200K ctx·$1.2/M in·$4/M out·Updated Jul 17, 2026Anonymous

GLM-5.1 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple…

200K ctx·$1.54/M in·$4.84/M out·Updated Jul 17, 2026Private

GLM 5.1 running in a Trusted Execution Environment (TEE). Hardware attestation evidence is available for independent verification of enclave identity and configuration.

200K ctx·$1.1/M in·$4.15/M out·Updated Jul 17, 2026E2EE

GLM-5.2 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple…

1M ctx·$1.4/M in·$4.4/M out·Updated Jul 17, 2026Private

GLM 5.2 running in a Trusted Execution Environment (TEE). Z.AI's flagship model for long-horizon tasks with enhanced reasoning and project-level engineering context, with hardware attestation…

524K ctx·$1.75/M in·$5.75/M out·Updated Jul 17, 2026E2EE

GLM-5V-Turbo is Z.ai's first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks with image, video, and text inputs.

200K ctx·$1.5/M in·$5/M out·Updated Jul 17, 2026Anonymous

Embedding model for semantic search and retrieval, served through Venice.

8K ctx·$0.15/M in·$0.6/M out·Updated Jul 17, 2026Private

Google's Lyria 3 Pro generates full-length, structured songs up to 3 minutes long from a single text prompt. Supports vocals, lyrics, and multi-language generation across genres.

ctx··$0.10/track·Updated Jul 17, 2026Anonymous

Image-to-video generation, served through Venice partner routing.

ctx··Varies·Updated Jul 17, 2026Private