Claude
Claude Fable 5
anthropic/claude-fable-5
Anonymous

Claude Fable 5 is Anthropic's most capable widely released model, designed for demanding reasoning and long-horizon agentic work. It features a 1M token context window, 128K max output tokens, always-on adaptive thinking, and strong multimodal capabilities.

Anonymous|1M context|$12.00/M input|$60.00/M output
OpenAI
GPT-5.6 Sol Pro
openai/gpt-5.6-sol-pro
Anonymous

GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

Anonymous|1M context|$2.50/M input|$12.50/M output
Venice
Venice Uncensored 1.2
venice/venice-uncensored-1.2
Private

Venice Uncensored 1.2 is designed for maximum creative freedom and authentic interaction. Built for open-ended exploration, roleplay, and unfiltered dialogue with improved capabilities over 1.1.

Private|128K context|$0.20/M input|$0.90/M output
Kimi
Kimi K3
moonshotai/kimi-k3
TEE

Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.

TEE|1M context|$3.75/M input|$18.75/M output
Claude
Claude Opus 4.8
anthropic/claude-opus-4.8
Anonymous

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports long-horizon agentic work, complex multi-step coding, and memory-driven tasks where coherence over extended sessions matters. It features a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.

Anonymous|1M context|$6.00/M input|$30.00/M output
Grok
Grok 4.5
x-ai/grok-4.5
Private

Grok 4.5 is xAI's intelligent coding model for agentic software engineering and workflow tasks, with function calling, structured outputs, and a 500K-token context window.

Private|500K context|$2.27/M input|$6.80/M output
Gemini
Gemini 3.5 Flash
google/gemini-3.5-flash
Anonymous

Gemini 3.5 Flash is a high speed, high value thinking model with 1M context, designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro level reasoning with substantially lower latency.

Anonymous|1M context|$1.55/M input|$9.45/M output
DeepSeek
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
Private

DeepSeek V4 Pro is a 1.6T-parameter Mixture-of-Experts model with 49B active parameters and a 1M-token context window. Built for advanced reasoning, coding, and long-horizon agentic workflows with a hybrid attention system for efficient long-context processing.

Private|1M context|$1.65/M input|$3.301/M output
Zhipu
GLM 5.2
z-ai/glm-5.2
E2EE

GLM 5.2 running in a Trusted Execution Environment (TEE). Z.AI's flagship model for long-horizon tasks with enhanced reasoning and project-level engineering context, with hardware attestation evidence available for independent verification.

E2EE|524K context|$1.75/M input|$5.75/M output
Qwen
Qwen 3.7 Max
qwen/qwen-3.7-max
Anonymous

Qwen 3.7 Max is the largest model in the Qwen 3.7 series, with deep thinking, function calling, prompt caching, and multimodal input support for images and video. It excels at programming, office and productivity tasks, and long-running autonomous agent workflows.

Anonymous|1M context|$2.70/M input|$8.05/M output
OpenAI
GPT-5.5
openai/gpt-5.5
Anonymous

GPT-5.5 is the latest frontier model in the GPT-5 series with a 1M+ context window, offering improved agentic and long context performance. It uses adaptive reasoning to dynamically allocate computation across tasks.

Anonymous|1M context|$6.25/M input|$37.50/M output
Minimax
MiniMax M3 Preview
minimax/minimax-m3-preview
Private

MiniMax-M3 preview is a 1.4T-parameter frontier model from MiniMax for coding, agentic workflows, and complex reasoning, served at fp8 with a 512K context window.

Private|524K context|$0.30/M input|$1.20/M output
AionLabs
Aion 3.0
aion-labs/aion-3.0
Anonymous

Aion 3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. Multiple specialized models collaborate on each response to produce stronger narrative structure and more compelling tension and conflict. It handles mature and darker themes with nuance and supports tool calling for richer interactive fiction.

Anonymous|128K context|$3.75/M input|$7.50/M output
AionLabs
Aion 3.0 Mini
aion-labs/aion-3.0-mini
Anonymous

Aion 3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. Multiple specialized models collaborate on each response to produce stronger narrative structure and more compelling tension and conflict at lower cost. It handles mature and darker themes with nuance and supports tool calling for richer interactive fiction.

Anonymous|128K context|$0.875/M input|$1.75/M output
Alibaba
Wan 2.7 Pro
alibaba/wan-2.7-pro
Anonymous

Wan 2.7 Pro is a text-to-image model. High-fidelity results with strong prompt adherence and fast turnaround.

Anonymous|Unknown context||$0.09/image
Alibaba
Wan 2.7
alibaba/wan-2.7-t2i
Anonymous

Wan 2.7 is a text-to-image model. High-fidelity results with strong prompt adherence and fast turnaround.

Anonymous|Unknown context||$0.04/image
Alibaba
Z-Image Turbo
alibaba/z-image-turbo
Private

Z-Image Turbo is a text-to-image model. High-fidelity results with strong prompt adherence and fast turnaround.

Private|Unknown context||$0.01/image
Claude
Claude Fable 5.1
anthropic/claude-fable-5.1
Anonymous

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work such as long code refactors, front-end development, and finance and analysis tasks. It features a 1M token context window, 128K max output tokens, always-on adaptive thinking, and strong multimodal capabilities, and tends to be more concise in its plans and summaries.

Anonymous|1M context|$12.00/M input|$60.00/M output
Claude
Claude Opus 4.5
anthropic/claude-opus-4.5
Anonymous

Claude Opus 4.5 is Anthropic's frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and reasoning benchmarks, and improved robustness to prompt injection.

Anonymous|198K context|$6.00/M input|$30.00/M output
Claude
Claude Opus 4.6
anthropic/claude-opus-4.6
Anonymous

Claude Opus 4.6 is Anthropic's most capable reasoning model, building on Opus 4.5 with enhanced performance across complex software engineering, agentic workflows, and long-horizon tasks. It features a 1M token context window, improved multimodal capabilities, and stronger robustness to prompt injection.

Anonymous|1M context|$6.00/M input|$30.00/M output
Claude
Claude Opus 4.7
anthropic/claude-opus-4.7
Anonymous

Claude Opus 4.7 is Anthropic's most capable generally available model for complex reasoning and agentic coding. It features a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.

Anonymous|1M context|$6.00/M input|$30.00/M output
Claude
Claude Opus 4.8 Fast
anthropic/claude-opus-4.8-fast
Anonymous

Claude Opus 4.8 (Fast) is a speed-optimized variant of Anthropic's most capable generally available Opus model, offering the same 1M token context window and strong performance across long-horizon agentic work and complex coding — with lower latency.

Anonymous|1M context|$12.00/M input|$60.00/M output
Claude
Claude Opus 5
anthropic/claude-opus-5
Anonymous

Claude Opus 5 is Anthropic's most capable model in the Opus family. It delivers major gains over Opus 4.8 in agentic coding, professional knowledge work, and long-horizon reasoning, with a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.

Anonymous|1M context|$6.00/M input|$30.00/M output
Claude
Claude Opus 5 Fast
anthropic/claude-opus-5-fast
Anonymous

Claude Opus 5 (Fast) is a speed-optimized variant of Anthropic's most capable Opus model, offering the same 1M token context window and strong performance across agentic coding, professional knowledge work, and long-horizon reasoning — with lower latency.

Anonymous|1M context|$12.00/M input|$60.00/M output
Claude
Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anonymous

Claude Sonnet 4.5 is Anthropic's balanced model offering strong performance on coding, reasoning, and general tasks with good speed and cost efficiency.

Anonymous|198K context|$3.75/M input|$18.75/M output
Claude
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anonymous

Claude Sonnet 4.6 is Anthropic's best combination of speed and intelligence, offering strong performance on coding, reasoning, and general tasks with excellent speed and cost efficiency. It features a 1M token context window and 64K max output tokens.

Anonymous|1M context|$3.60/M input|$18.00/M output
Claude
Claude Sonnet 5
anthropic/claude-sonnet-5
Anonymous

Claude Sonnet 5 is Anthropic's latest Sonnet model, substantially improving on Sonnet 4.6 in coding and agentic work and reaching near-Opus quality on many tasks. It features a 1M token context window, adaptive thinking, and strong document and vision understanding.

Anonymous|1M context|$3.00/M input|$15.00/M output
BAAI
BGE-EN-ICL
baai/bge-en-icl
Private

Embedding model for semantic search and retrieval.

Private|8K context|$0.0125/M input|$0.0125/M output
BAAI
BGE-M3
baai/bge-m3
Private

Embedding model for semantic search and retrieval.

Private|8K context|$0.15/M input|$0.60/M output
Flux
Flux 2 Max
black-forest-labs/flux-2-max
Anonymous

Flux 2 Max is a text-to-image model. Photorealistic, high-detail output with excellent typography and layout.

Anonymous|Unknown context||$0.09/image
Flux
Flux 2 Pro
black-forest-labs/flux-2-pro
Anonymous

Flux 2 Pro is a text-to-image model. Photorealistic, high-detail output with excellent typography and layout.

Anonymous|Unknown context||$0.03/image
ByteDance
Seedream V4.5
bytedance/seedream-v4.5
Anonymous

Seedream V4.5 is a text-to-image model. Vivid color, sharp detail, and coherent multi-subject scenes.

Anonymous|Unknown context||$0.05/image
ByteDance
Seedream V5 Lite
bytedance/seedream-v5-lite
Anonymous

Seedream V5 Lite is a text-to-image model. Vivid color, sharp detail, and coherent multi-subject scenes.

Anonymous|Unknown context||$0.05/image
DeepSeek
DeepSeek V3.2
deepseek/deepseek-v3.2
E2EE

DeepSeek-V3.2 is an efficient large language model with DeepSeek Sparse Attention (DSA) for long contexts. It features strong reasoning and tool-use skills, achieving top results on the 2025 IMO and IOI.

E2EE|160K context|$0.33/M input|$0.48/M output
DeepSeek
DeepSeek V4 Flash 0423
deepseek/deepseek-v4-flash-0423
Private

DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.

Private|1M context|$0.138/M input|$0.275/M output
DeepSeek
DeepSeek V4 Flash 0731 Fast
deepseek/deepseek-v4-flash-0731-fast
Private

DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.

Private|1M context|$0.35/M input|$0.70/M output
DeepSeek
DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813
Private

DeepSeek V4 Pro is a 1.6T-parameter Mixture-of-Experts model with 49B active parameters and a 1M-token context window. Built for advanced reasoning, coding, and long-horizon agentic workflows with a hybrid attention system for efficient long-context processing.

Private|1M context|$1.65/M input|$4.95/M output
ElevenLabs
ElevenLabs Turbo v2.5
elevenlabs/elevenlabs-turbo-v2.5
Anonymous

ElevenLabs Turbo v2.5 is a text-to-speech model. Natural, expressive speech with clear articulation.

Anonymous|Unknown context|$62.50/1M chars|
Gemini
Gemini 3 Flash Preview
google/gemini-3-flash-preview
Anonymous

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro level reasoning with substantially lower latency.

Anonymous|256K context|$0.70/M input|$3.75/M output
Gemini
Gemini 3.1 Flash TTS
google/gemini-3.1-flash-tts
Anonymous

Gemini 3.1 Flash TTS is a text-to-speech model. Natural, expressive speech with clear articulation.

Anonymous|Unknown context|$187.50/1M chars|
Gemini
Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
Anonymous

Gemini 3.1 Pro is the latest evolution of Google flagship frontier model with 1M context, advancing high-precision multimodal reasoning across text, image, and code.

Anonymous|1M context|$2.50/M input|$15.00/M output
Gemini
Gemini 3.5 Flash-Lite
google/gemini-3.5-flash-lite
Anonymous

Gemini 3.5 Flash-Lite is the fastest, most cost-efficient Gemini 3.5 model with 1M context, ideal for everyday questions, summarization, and lightweight coding tasks.

Anonymous|1M context|$0.375/M input|$3.125/M output
Gemini
Gemini 3.6 Flash
google/gemini-3.6-flash
Anonymous

Gemini 3.6 Flash is a high speed, high value thinking model with 1M context, designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro level reasoning with substantially lower latency.

Anonymous|1M context|$0.9375/M input|$4.6875/M output
Gemini
Gemini 3.7 Flash
google/gemini-3.7-flash
Anonymous

Gemini 3.7 Flash is Google's most capable Flash model, built for complex coding, agentic workflows, and reliable multi-step execution, with 1M context and tunable thinking.

Anonymous|1M context|$0.9375/M input|$4.6875/M output
Gemini
Gemini 3.8 Flash
google/gemini-3.8-flash
Anonymous

Gemini 3.8 Flash is Google's most capable Flash model, built for complex coding, agentic workflows, and reliable multi-step execution, with 1M context and tunable thinking.

Anonymous|1M context|$0.9375/M input|$4.6875/M output
Gemini
Gemini Embedding 2 Preview
google/gemini-embedding-2-preview
Anonymous

Embedding model for semantic search and retrieval.

Anonymous|2K context|$0.25/M input|$0.25/M output
Gemma
Gemma 4 26B A4B Uncensored
google/gemma-4-26b-a4b-uncensored
E2EE

Gemma 4 26B A4B Uncensored running in a Trusted Execution Environment (TEE). An uncensored variant of Google's Gemma 4 MoE model with 25.2B total / 3.8B active parameters, supporting multimodal input across text and images, with hardware attestation evidence available for independent verification.

E2EE|64K context|$0.19/M input|$0.88/M output
Gemma
Google Gemma 4 31B Instruct
google/gemma-4-31b-instruct
TEE

Gemma 4 31B is a dense model from Google DeepMind with 31B parameters, delivering frontier-level reasoning performance. It handles text, image, and video input, supports 256K context, function calling, and configurable thinking modes.

TEE|256K context|$0.12/M input|$0.36/M output
Gemma
Gemma 4 Uncensored
google/gemma-4-uncensored
Private

Gemma 4 Uncensored is an uncensored variant of Google Gemma 4 26B, a Mixture-of-Experts model with 26B total parameters and only 4B active per token. Fine-tuned for uncensored chat without content filtering, it supports 256K context, coding, and general-purpose conversation.

Private|256K context|$0.1625/M input|$0.50/M output
Gemma
Google Gemma 3 27B Instruct
google/google-gemma-3-27b-instruct
Private

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to Gemma 2.

Private|198K context|$0.12/M input|$0.20/M output
Gemma
Google Gemma 4 26B A4B Instruct
google/google-gemma-4-26b-a4b-instruct
Private

Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.

Private|256K context|$0.13/M input|$0.40/M output
NanoBanana
Nano Banana 2 Lite
google/nano-banana-2-lite
Anonymous

Nano Banana 2 Lite is a text-to-image model. Creative, high-quality results with strong instruction following and clean edits.

Anonymous|Unknown context||$0.06/image
Ideogram
Ideogram V4
ideogram/ideogram-v4
Anonymous

Ideogram V4 is a text-to-image model. Best-in-class in-image text and typography.

Anonymous|Unknown context||$0.06/image
Inception
Mercury 2
inception/mercury-2
Anonymous

Mercury 2 is a diffusion-based reasoning LLM from Inception, delivering over 1,000 tokens per second — 5x faster than leading speed-optimized models — with strong reasoning, tool use, and structured output capabilities.

Anonymous|128K context|$0.3125/M input|$0.9375/M output
Krea
Krea v2 Large
krea/krea-v2-large
Anonymous

Krea v2 Large is a text-to-image model. Aesthetic-focused, with a distinctive and stylized look.

Anonymous|Unknown context||$0.07/image
Krea
Krea v2 Medium
krea/krea-v2-medium
Anonymous

Krea v2 Medium is a text-to-image model. Aesthetic-focused, with a distinctive and stylized look.

Anonymous|Unknown context||$0.04/image
Luma
Luma Uni-1
luma/luma-uni-1
Anonymous

Luma Uni-1 is a text-to-image model. Cinematic lighting and composition with photoreal detail.

Anonymous|Unknown context||$0.05/image
Luma
Luma Uni-1 Max
luma/luma-uni-1-max
Anonymous

Luma Uni-1 Max is a text-to-image model. Cinematic lighting and composition with photoreal detail.

Anonymous|Unknown context||$0.12/image
Meta
Hermes 3 Llama 3.1 405b
meta-llama/hermes-3-llama-3.1-405b
Private

Hermes 3 405B is a frontier level, full parameter finetune of the Llama-3.1 405B foundation model, focused on aligning LLMs to the user, with powerful steering capabilities and control given to the end user.

Private|128K context|$1.10/M input|$3.00/M output
Meta
Llama 3.2 3B
meta-llama/llama-3.2-3b
Private

Llama 3.2 3B is a text model.

Private|128K context|$0.15/M input|$0.60/M output
Meta
Llama 3.3 70B
meta-llama/llama-3.3-70b
TEE

Llama 3.3 70B is a text model.

TEE|128K context|$0.70/M input|$2.80/M output
Azure
Multilingual E5 Large Instruct
microsoft/multilingual-e5-large-instruct
Private

Embedding model for semantic search and retrieval.

Private|512 context|$0.0125/M input|$0.0125/M output
Minimax
MiniMax M2.5
minimax/minimax-m2.5
Private

MiniMax-M2.5 is a state-of-the-art large language model optimized for coding, agentic workflows, and modern application development with enhanced reasoning capabilities.

Private|198K context|$0.27/M input|$0.95/M output
Minimax
MiniMax M2.7
minimax/minimax-m2.7
Private

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity with advanced agentic capabilities through multi-agent collaboration.

Private|198K context|$0.375/M input|$1.50/M output
Minimax
MiniMax Speech-02 HD
minimax/minimax-speech-02-hd
Anonymous

Clone your voice from a short recording and generate natural speech in it across 30+ languages.

Anonymous|Unknown context|$125.00/1M chars|
Mistral
Mistral Small 3.2 24B Instruct
mistralai/mistral-small-3.2-24b-instruct
Private

Mistral Small 3.2 is a 24B parameter model optimized for efficiency and performance. Ideal for general-purpose tasks with balanced speed and capability.

Private|256K context|$0.0938/M input|$0.25/M output
Mistral
Mistral Small 4
mistralai/mistral-small-4
Private

Mistral Small 4 unifies instruction following, reasoning, coding, and vision in a single 119B MoE model with 256K context and configurable reasoning effort.

Private|256K context|$0.1875/M input|$0.75/M output
Kimi
Kimi K2.5
moonshotai/kimi-k2.5
Private

Kimi K2.5 is Moonshot AIs most advanced open reasoning model, featuring trillion-parameter Mixture-of-Experts architecture with 32B active parameters and 256K context windows.

Private|256K context|$0.56/M input|$3.50/M output
Kimi
Kimi K2.6
moonshotai/kimi-k2.6
E2EE

Kimi K2.6 is an open-source, native multimodal agentic model from Moonshot AI with 1T total parameters and 32B active parameters. It excels at long-horizon coding, coding-driven design, agent swarm orchestration, and proactive autonomous execution with 256K context windows.

E2EE|256K context|$0.75/M input|$3.50/M output
Kimi
Kimi K2.7 Code
moonshotai/kimi-k2.7-code
Private

Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built on Kimi K2.6, with 1T total parameters and 32B active parameters. It always operates in thinking mode, supports text and image input, and targets long-horizon software engineering, agentic task decomposition, and multi-turn coding workflows with 256K context.

Private|256K context|$0.75/M input|$3.50/M output
Kimi
Kimi K3 Fast
moonshotai/kimi-k3-fast
Private

Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.

Private|1M context|$4.50/M input|$22.50/M output
Nvidia
Nemotron Embed VL 1B v2
nvidia/nemotron-embed-vl-1b-v2
Private

Embedding model for semantic search and retrieval.

Private|33K context|$0.0125/M input|$0.0125/M output
Nvidia
NVIDIA Nemotron 3 Nano 30B
nvidia/nvidia-nemotron-3-nano-30b
Private

NVIDIA Nemotron 3 Nano 30B is a compact and efficient language model from NVIDIA, optimized for fast inference while maintaining strong performance across diverse tasks.

Private|128K context|$0.075/M input|$0.30/M output
Nvidia
NVIDIA Nemotron 3 Ultra
nvidia/nvidia-nemotron-3-ultra
Private

NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

Private|256K context|$0.625/M input|$3.125/M output
OpenAI
GPT-4o
openai/gpt-4o
Anonymous

OpenAI's multimodal flagship model with vision capabilities, strong reasoning, and broad knowledge. Popular for its balanced performance across tasks. Version: 2024-11-20.

Anonymous|128K context|$3.125/M input|$12.50/M output
OpenAI
GPT-4o Mini
openai/gpt-4o-mini
Anonymous

OpenAI's cost-efficient small model that delivers GPT-4 level intelligence at a fraction of the cost. Ideal for high-volume applications requiring strong reasoning. Version: 2024-07-18.

Anonymous|128K context|$0.1875/M input|$0.75/M output
OpenAI
GPT-5.2
openai/gpt-5.2
Anonymous

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context performance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly to simple queries while spending more depth on complex tasks.

Anonymous|256K context|$2.19/M input|$17.50/M output
OpenAI
GPT-5.2 Codex
openai/gpt-5.2-codex
Anonymous

GPT-5.2 Codex is OpenAI specialized coding model built on GPT-5.2, optimized for advanced software development, code generation, and technical problem-solving.

Anonymous|256K context|$2.19/M input|$17.50/M output
OpenAI
GPT-5.3 Codex
openai/gpt-5.3-codex
Anonymous

GPT-5.3 Codex is OpenAI specialized coding model built on GPT-5.3, optimized for advanced software development, code generation, and technical problem-solving.

Anonymous|400K context|$2.19/M input|$17.50/M output
OpenAI
GPT-5.4
openai/gpt-5.4
Anonymous

GPT-5.4 is the latest frontier model in the GPT-5 series with a 1M+ context window, offering improved agentic and long context performance. It uses adaptive reasoning to dynamically allocate computation across tasks.

Anonymous|1M context|$3.13/M input|$18.80/M output
OpenAI
GPT-5.4 Mini
openai/gpt-5.4-mini
Anonymous

GPT-5.4 Mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use.

Anonymous|400K context|$0.9375/M input|$5.625/M output
OpenAI
GPT-5.4 Pro
openai/gpt-5.4-pro
Anonymous

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning for complex, high-stakes tasks. It provides a 1M+ token context window (922K input, 128K output) and supports text and image inputs.

Anonymous|1M context|$37.50/M input|$225.00/M output
OpenAI
GPT-5.5 Pro
openai/gpt-5.5-pro
Anonymous

GPT-5.5 Pro is OpenAI's most advanced model, building on GPT-5.5's unified architecture with enhanced reasoning for complex, high-stakes tasks. It provides a 1M+ token context window (922K input, 128K output) and supports text and image inputs.

Anonymous|1M context|$37.50/M input|$225.00/M output
OpenAI
GPT-5.6 Terra
openai/gpt-5.6-terra
Anonymous

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks where capability and cost need to be balanced.

Anonymous|1M context|$2.50/M input|$15.00/M output
OpenAI
GPT-5.6 Terra Pro
openai/gpt-5.6-terra-pro
Anonymous

GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

Anonymous|1M context|$2.50/M input|$15.00/M output
OpenAI
GPT-6 Astra
openai/gpt-6-astra
Anonymous

GPT-6 Astra is OpenAI's most capable model, built for the hardest end-to-end work. It is suited for complex reasoning, coding, computer use, research, and document creation, with a 1.05M token context window (922K input, 128K output) and support for text and image inputs.

Anonymous|1M context|$10.00/M input|$50.00/M output
OpenAI
GPT-6 Astra Pro
openai/gpt-6-astra-pro
Anonymous

GPT-6 Astra with pro reasoning mode for difficult tasks that benefit from more model work. Supports text and image inputs with a 1.05M token context window. Pro mode can use more tokens and take longer than standard Astra.

Anonymous|1M context|$12.50/M input|$62.50/M output
OpenAI
GPT Image 1.5
openai/gpt-image-1.5
Anonymous

GPT Image 1.5 is a text-to-image model. Instruction-following generation with broad world knowledge and reliable text rendering.

Anonymous|Unknown context||$0.26/image
OpenAI
OpenAI GPT OSS 120B
openai/gpt-oss-120b
E2EE

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation

E2EE|128K context|$0.07/M input|$0.30/M output
OpenAI
Text Embedding 3 Large
openai/text-embedding-3-large
Anonymous

Embedding model for semantic search and retrieval.

Anonymous|8K context|$0.1625/M input|$0.1625/M output
OpenAI
Text Embedding 3 Small
openai/text-embedding-3-small
Anonymous

Embedding model for semantic search and retrieval.

Anonymous|8K context|$0.025/M input|$0.025/M output
Qwen
Qwen 2.5 7B
qwen/qwen-2.5-7b
E2EE

Qwen 2.5 7B Instruct running in a Trusted Execution Environment (TEE). A compact model with strong coding, math, and multilingual capabilities supporting 29+ languages, with hardware attestation evidence available for independent verification.

E2EE|32K context|$0.05/M input|$0.13/M output
Qwen
Qwen 3 235B A22B Instruct 2507
qwen/qwen-3-235b-a22b-instruct-2507
Private

Built for in-depth research and handling long, complex documents. Ideal for technical work, multimodal input, and high-precision tasks.

Private|128K context|$0.15/M input|$0.75/M output
Qwen
Qwen 3 235B A22B Thinking 2507
qwen/qwen-3-235b-a22b-thinking-2507
Private

Built for in-depth research and handling long, complex documents. Ideal for technical work, multimodal input, and high-precision tasks.

Private|128K context|$0.45/M input|$3.50/M output
Qwen
Qwen 3 Coder 480B Turbo
qwen/qwen-3-coder-480b-turbo
Private

Turbo variant of Qwen3 Coder 480B, optimized for faster inference on code tasks.

Private|256K context|$0.35/M input|$1.50/M output
Qwen
Qwen 3 Next 80b
qwen/qwen-3-next-80b
Private

Optimized for speed and efficiency.

Private|256K context|$0.35/M input|$1.90/M output
Qwen
Qwen 3 TTS 0.6B
qwen/qwen-3-tts-0.6b
Private

Qwen 3 TTS 0.6B is a text-to-speech model. Natural, expressive speech with clear articulation.

Private|Unknown context|$87.50/1M chars|
Qwen
Qwen 3 TTS 1.7B
qwen/qwen-3-tts-1.7b
Private

Qwen 3 TTS 1.7B is a text-to-speech model. Natural, expressive speech with clear articulation.

Private|Unknown context|$112.50/1M chars|
Qwen
Qwen 3.5 35B A3B
qwen/qwen-3.5-35b-a3b
Private

Qwen 3.5 35B A3B is a highly efficient MoE model with 35B total parameters and only 3B active parameters. It surpasses the larger Qwen3-235B-A22B while being 6.7x smaller, excelling at reasoning, coding, and general knowledge tasks.

Private|256K context|$0.3125/M input|$1.25/M output
Qwen
Qwen 3.5 397B
qwen/qwen-3.5-397b
Private

Qwen 3.5 is Alibaba flagship reasoning model featuring a 397B parameter Mixture-of-Experts architecture with 17B active parameters. It excels at complex reasoning, coding, and general knowledge tasks.

Private|128K context|$0.75/M input|$4.50/M output
Qwen
Qwen 3.5 9B
qwen/qwen-3.5-9b
Private

A 9B dense model with 262K native context window (extendable to 1M). Features Gated DeltaNet hybrid attention architecture for efficient long-context processing. Supports 201 languages, thinking/reasoning mode, and function calling.

Private|256K context|$0.10/M input|$0.15/M output
Qwen
Qwen 3.6 27B
qwen/qwen-3.6-27b
Private

The Qwen 3.6 27B native vision-language dense model builds upon the 3.5-27B version, with key improvements in agentic coding capabilities and enhanced STEM reasoning and inference skills. In the vision modality, it demonstrates significant advances in spatial intelligence, object localization, and detection, while video understanding, document OCR, and visual agent capabilities continue to improve steadily.

Private|256K context|$0.325/M input|$3.25/M output
Qwen
Qwen 3.6 35B A3B
qwen/qwen-3.6-35b-a3b
Private

Qwen 3.6 35B A3B is a fast mixture-of-experts model with 35B total parameters and ~3B active per token. Strong at agentic coding, STEM reasoning, and tool use, with a native 256K context window.

Private|256K context|$0.10/M input|$1.00/M output
Qwen
Qwen 3.6 Plus Uncensored
qwen/qwen-3.6-plus-uncensored
Anonymous

Qwen 3.6 Plus Uncensored is Alibaba's latest flagship reasoning model with exceptional performance across coding, reasoning, and general knowledge tasks. Features mixed reasoning, function calling, and multimodal input support.

Anonymous|1M context|$0.625/M input|$3.75/M output
Qwen
Qwen 3.7 Plus
qwen/qwen-3.7-plus
Anonymous

Qwen 3.7 Plus is Alibaba's latest flagship reasoning model with exceptional performance across coding, reasoning, and general knowledge tasks. Features mixed reasoning, function calling, and multimodal input support.

Anonymous|1M context|$0.50/M input|$2.00/M output
Qwen
Qwen 3.8 2.4T
qwen/qwen-3.8-2.4t
Private

Qwen 3.8 2.4T is Alibaba's open-weight 2.4-trillion-parameter MoE model (95B active), with major gains in software engineering, research, and long-horizon agentic tasks. It is text-only, requires thinking mode, and supports a 262K-token context window.

Private|262K context|$2.50/M input|$7.50/M output
Qwen
Qwen 3.8 27B
qwen/qwen-3.8-27b
E2EE

Qwen 3.8 27B is a native vision-language dense model with 27B parameters. It improves coding, professional work, research, and long-horizon agentic tasks, with flexible thinking control and image and video understanding. It supports a native 262K-token context window.

E2EE|262K context|$0.45/M input|$3.20/M output
Qwen
Qwen 3.8 Max
qwen/qwen-3.8-max
Anonymous

Qwen 3.8 Max is Alibaba's flagship 2.4-trillion-parameter MoE model, with major gains over Qwen 3.7 Max in software engineering and office-productivity workflows and strong long-horizon, multi-agent performance. It accepts both text and vision-language input (images and video), operates in thinking mode only, and supports a 1M-token context window.

Anonymous|1M context|$2.50/M input|$7.50/M output
Qwen
Qwen Image
qwen/qwen-image
Anonymous

Qwen Image is a text-to-image model. Flexible across a wide range of prompts and edits.

Anonymous|Unknown context||$0.03/image
Qwen
Qwen Image 2
qwen/qwen-image-2
Anonymous

Qwen Image 2 is a text-to-image model. Flexible across a wide range of prompts and edits.

Anonymous|Unknown context||$0.05/image
Qwen
Qwen Image 2 Pro
qwen/qwen-image-2-pro
Anonymous

Qwen Image 2 Pro is a text-to-image model. Flexible across a wide range of prompts and edits.

Anonymous|Unknown context||$0.10/image
Qwen
Qwen Image 3
qwen/qwen-image-3
Anonymous

Qwen Image 3 is a text-to-image model. Flexible across a wide range of prompts and edits.

Anonymous|Unknown context||$0.04/image
Qwen
Qwen3 Embedding 0.6B
qwen/qwen3-embedding-0.6b
Private

Embedding model for semantic search and retrieval.

Private|33K context|$0.0125/M input|$0.0125/M output
Qwen
Qwen3 Embedding 8B
qwen/qwen3-embedding-8b
Private

Embedding model for semantic search and retrieval.

Private|33K context|$0.0125/M input|$0.0125/M output
Qwen
Qwen3 VL 235B
qwen/qwen3-vl-235b
Private

Qwen3-VL 235B vision-language model with MoE architecture. The most powerful VL model in the Qwen series with superior visual perception, OCR, and multimodal reasoning.

Private|128K context|$0.21/M input|$1.90/M output
Recraft
Recraft V4
recraft/recraft-v4
Anonymous

Recraft V4 is a text-to-image model. Brand- and vector-friendly, with precise control over style.

Anonymous|Unknown context||$0.05/image
Recraft
Recraft V4 Pro
recraft/recraft-v4-pro
Anonymous

Recraft V4 Pro is a text-to-image model. Brand- and vector-friendly, with precise control over style.

Anonymous|Unknown context||$0.29/image
Stability
Lustify SDXL
stabilityai/lustify-sdxl
Private

Lustify SDXL is a text-to-image model. Open-weight, with fine-grained control over style and composition.

Private|Unknown context||$0.01/image
Stability
Venice SD35
stabilityai/venice-sd35
Private

Venice SD35 is a text-to-image model. Open-weight, with fine-grained control over style and composition.

Private|Unknown context||$0.01/image
Hunyuan
Hunyuan Image 3.0
tencent/hunyuan-image-3.0
Private

Hunyuan Image 3.0 is a text-to-image model. Smooth motion with strong scene and subject consistency.

Private|Unknown context||$0.09/image
Venice
Inkling
thinking-machines/inkling
Private

Inkling is a general-purpose multimodal model from Thinking Machines Lab that accepts text, image, and audio inputs and generates text. It is a 66-layer sparse MoE (975B total / 41B active) with hybrid local/global attention, 512K context, and variable thinking effort — suited for chat, coding, tool use, and agentic workflows. Video input is not supported on Venice.

Private|524K context|$1.25/M input|$5.0625/M output
Venice
Anime (WAI)
venice/anime-wai
Private

Anime (WAI) is a text-to-image model.

Private|Unknown context||$0.01/image
Venice
Background Remover
venice/background-remover
Anonymous

Background Remover is a text-to-image model.

Anonymous|Unknown context||$0.03/image
Venice
Chatterbox HD (Resemble AI)
venice/chatterbox-hd-resemble-ai
Private

Chatterbox HD (Resemble AI) is a text-to-speech model. Natural, expressive speech with clear articulation.

Private|Unknown context|$50.00/1M chars|
Venice
Chroma
venice/chroma
Private

Chroma is a text-to-image model.

Private|Unknown context||$0.01/image
Venice
Gradium TTS
venice/gradium-tts
Anonymous

Gradium TTS is a text-to-speech model. Natural, expressive speech with clear articulation.

Anonymous|Unknown context|$47.50/1M chars|
Venice
ImagineArt 1.5 Pro
venice/imagineart-1.5-pro
Anonymous

ImagineArt 1.5 Pro is a text-to-image model.

Anonymous|Unknown context||$0.06/image
Venice
Inworld TTS-1.5 Max
venice/inworld-tts-1.5-max
Anonymous

Inworld TTS-1.5 Max is a text-to-speech model. Natural, expressive speech with clear articulation.

Anonymous|Unknown context|$12.50/1M chars|
Venice
Kokoro Text to Speech
venice/kokoro-text-to-speech
Private

Kokoro Text to Speech is a text-to-speech model. Natural, expressive speech with clear articulation.

Private|Unknown context|$3.50/1M chars|
Venice
Lustify v7
venice/lustify-v7
Private

Lustify v7 is a text-to-image model.

Private|Unknown context||$0.01/image
Venice
Lustify v8
venice/lustify-v8
Private

Lustify v8 is a text-to-image model.

Private|Unknown context||$0.01/image
Venice
Muse Image
venice/muse-image
Anonymous

Muse Image is a text-to-image model.

Anonymous|Unknown context||$0.02/image
Venice
Orpheus TTS
venice/orpheus-tts
Private

Orpheus TTS is a text-to-speech model. Natural, expressive speech with clear articulation.

Private|Unknown context|$62.50/1M chars|
Venice
Seed 2.1 Turbo
venice/seed-2.1-turbo
Anonymous

Seed 2.1 Turbo (Dola-Seed-2.1) is ByteDance’s next-generation multimodal model for the coding and agent era, with engineering-grade code delivery, long-horizon agent execution, and upgraded GUI and video understanding. Supports text, image, and video inputs with a 256K context window.

Anonymous|256K context|$0.625/M input|$3.125/M output
Venice
Venice Role Play Uncensored
venice/venice-role-play-uncensored
Private

Optimized for creative roleplay scenarios with maximum freedom. Designed for immersive storytelling, character interactions, and open-ended creative writing.

Private|128K context|$0.50/M input|$2.00/M output
Grok
Grok 4.20
x-ai/grok-4.20
Private

Grok 4.20 is xAI's latest multimodal reasoning model with strong tool use, structured output support, and a 2M-token context window.

Private|2M context|$1.42/M input|$2.83/M output
Grok
Grok 4.20 Multi-Agent
x-ai/grok-4.20-multi-agent
Private

Grok 4.20 Multi-Agent is a variant of xAI Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information across complex tasks.

Private|2M context|$1.42/M input|$2.83/M output
Grok
Grok 4.3
x-ai/grok-4.3
Private

Grok 4.3 is xAI's most intelligent and fastest reasoning model with function calling, structured outputs, and a 1M-token context window. Suited for agentic workflows, instruction-following tasks, and applications requiring high factual accuracy.

Private|1M context|$1.42/M input|$2.83/M output
Grok
Grok 4.6
x-ai/grok-4.6
Private

Grok 4.6 is xAI's multimodal chat and reasoning model with function calling, structured outputs, adjustable reasoning effort (low/medium/high/xhigh), and a 500K-token context window.

Private|500K context|$2.27/M input|$6.80/M output
Grok
Grok Build 0.1
x-ai/grok-build-0.1
Private

xAI's fast coding model trained specifically for agentic coding, currently in early access.

Private|256K context|$1.00/M input|$2.00/M output
Grok
xAI TTS v1
x-ai/xai-tts-v1
Anonymous

xAI TTS v1 is a text-to-speech model. Natural, expressive speech with clear articulation.

Anonymous|Unknown context|$18.75/1M chars|
Zhipu
GLM 4.6
z-ai/glm-4.6
Private

GLM-4.6 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.

Private|198K context|$0.43/M input|$1.75/M output
Zhipu
GLM 4.7
z-ai/glm-4.7
Private

GLM-4.7 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.

Private|198K context|$0.55/M input|$2.65/M output
Zhipu
GLM 4.7 Flash Heretic
z-ai/glm-4.7-flash-heretic
Private

GLM-4.7-Flash-Heretic is an uncensored experimental variant of GLM-4.7-Flash, optimized for creative freedom and unfiltered dialogue with fast inference speed.

Private|200K context|$0.07/M input|$0.40/M output
Zhipu
GLM 5
z-ai/glm-5
Private

GLM-5 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple languages. Supports large context windows for processing extensive text and detailed analysis.

Private|198K context|$1.00/M input|$3.20/M output
Zhipu
GLM 5 Turbo
z-ai/glm-5-turbo
Anonymous

GLM-5 Turbo is a fast inference model from Z.ai tuned for strong performance in agent-driven environments and production coding workflows.

Anonymous|200K context|$1.20/M input|$4.00/M output
Zhipu
GLM 5.1
z-ai/glm-5.1
Private

GLM-5.1 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple languages. Supports large context windows for processing extensive text and detailed analysis with fast inference speed.

Private|200K context|$1.54/M input|$4.84/M output
Zhipu
GLM 5.3
z-ai/glm-5.3
E2EE

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

E2EE|1M context|$1.75/M input|$5.50/M output
Zhipu
GLM 5.3 Flash
z-ai/glm-5.3-flash
E2EE

GLM-5.3 Flash is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.

E2EE|1M context|$0.15/M input|$0.50/M output
Zhipu
GLM 5V Turbo
z-ai/glm-5v-turbo
Anonymous

GLM-5V-Turbo is Z.ai's first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks with image, video, and text inputs.

Anonymous|200K context|$1.50/M input|$5.00/M output
Claude
claude haiku 4 5
anthropic/claude-haiku-4-5
Anonymous

The next generation of Anthropic's fastest and most cost-effective model, optimal for use cases where speed and affordability matter.

Anonymous|200K context|$1.00/M input|$5.00/M output
DeepInfra
Seed 1.8
bytedance/seed-1.8
Anonymous

Optimized specifically for multimodal agent scenarios. It features enhanced agent capabilities, upgraded multimodal comprehension, and more flexible context management.

Anonymous|256K context|$0.25/M input|$2.00/M output
DeepInfra
Seed 2.0 code
bytedance/seed-2.0-code
Anonymous

A coding model optimized for real-world development environments, with reliable tool use in common IDEs such as Claude Code. It delivers strong front-end performance and supports Skills.

Anonymous|256K context|$0.50/M input|$3.00/M output
DeepInfra
Seed 2.0 mini
bytedance/seed-2.0-mini
Anonymous

Built for low-latency, high-concurrency, cost-sensitive use cases, with flexible deployment, four-tier thinking, and multimodal

Anonymous|256K context|$0.10/M input|$0.40/M output
DeepInfra
Seed 2.0 pro
bytedance/seed-2.0-pro
Anonymous

Built for the Agent era, it delivers stable performance in complex reasoning and long-horizon tasks, including multi-step planning, visual-text reasoning, video understanding, and advanced analysis.

Anonymous|256K context|$0.50/M input|$3.00/M output
DeepSeek
DeepSeek R1 0528
deepseek/deepseek-r1-0528
Private

The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528.

Private|164K context|$0.50/M input|$2.15/M output
DeepSeek
DeepSeek V3
deepseek/deepseek-v3
Private

DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2.

Private|164K context|$0.32/M input|$0.89/M output
DeepSeek
DeepSeek V3 0324
deepseek/deepseek-v3-0324
Private

DeepSeek-V3-0324, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token, an improved iteration over DeepSeek-V3.

Private|164K context|$0.24/M input|$0.90/M output
DeepSeek
DeepSeek V3.1
deepseek/deepseek-v3.1
Private

DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report. We have expanded our dataset by collecting additional long documents and substantially extending both training phases. The 32K extension phase has been increased 10-fold to 630B tokens, while the 128K extension phase has been extended by 3.3x to 209B tokens. Additionally, DeepSeek-V3.1 is trained using the UE8M0 FP8 scale data format to ensure compatibility with microscaling data formats.

Private|164K context|$0.25/M input|$0.95/M output
DeepSeek
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
E2EE

DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks.

E2EE|1M context|$0.09/M input|$0.18/M output
DeepSeek
DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
Private

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.

Private|1M context|$0.06/M input|$0.18/M output
DeepSeek
DeepSeek V4 Flash Vision Exp
deepseek/deepseek-v4-flash-vision-exp
Private

DeepSeek-V4-Flash-Vision-Exp is DeepSeek's experimental multimodal model in the V4-Flash family, adding visual understanding to the V4-Flash architecture. It serves a 1M-token (1,048,576) context window and supports image input with visual grounding, tool calling, structured/JSON output, and configurable reasoning effort (low/high/max, or disabled).

Private|1M context|$0.44/M input|$1.32/M output
Gemini
gemini 2.5 flash
google/gemini-2.5-flash
Anonymous

Gemini 2.5 Flash is Google's latest thinking model, designed to tackle increasingly complex problems. It's capable of reasoning through their thoughts before responding, resulting in enhanced performance and improved accuracy. Gemini 2.5 Flash: best for balancing reasoning and speed.

Anonymous|1M context|$0.30/M input|$2.50/M output
Gemini
gemini 2.5 pro
google/gemini-2.5-pro
Anonymous

Gemini 2.5 Pro is Google's the most advanced thinking model, designed to tackle increasingly complex problems. Gemini 2.5 Pro leads common benchmarks by meaningful margins and showcases strong reasoning and code capabilities. Gemini 2.5 models are thinking models, capable of reasoning through their thoughts before responding, resulting in enhanced performance and improved accuracy. The Gemini 2.5 Pro model is now available on DeepInfra.

Anonymous|1M context|$1.25/M input|$10.00/M output
Gemini
gemini 3.1 flash lite
google/gemini-3.1-flash-lite
Anonymous

Bring any idea to life with state-of-the-art reasoning to help you learn, build, and plan anything. Best for high-volume tasks that need efficiency and intelligence.

Anonymous|1M context|$0.25/M input|$1.50/M output
Gemini
gemini 3.1 pro
google/gemini-3.1-pro
Anonymous

Bring any idea to life with state-of-the-art reasoning to help you learn, build, and plan anything. Best for complex tasks and bringing creative concepts to life.

Anonymous|1M context|$2.00/M input|$12.00/M output
Gemma
gemma 3 12b it
google/gemma-3-12b-it
Private

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3-12B is Google's latest open source model, successor to Gemma 2

Private|131K context|$0.05/M input|$0.15/M output
Gemma
gemma 3 27b it
google/gemma-3-27b-it
Private

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to Gemma 2

Private|131K context|$0.08/M input|$0.16/M output
Gemma
gemma 3 4b it
google/gemma-3-4b-it
Private

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3-12B is Google's latest open source model, successor to Gemma 2

Private|131K context|$0.05/M input|$0.10/M output
Gemma
gemma 4 26B A4B it
google/gemma-4-26b-a4b-it
Private

Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

Private|262K context|$0.07/M input|$0.34/M output
Gemma
gemma 4 31B it turbo
google/gemma-4-31b-it-turbo
Private

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

Private|262K context|$0.09/M input|$0.34/M output
Gemma
gemma 4 31B it Ultra
google/gemma-4-31b-it-ultra
Anonymous

Ultra speed version of gemma-4-31B-it

Anonymous|131K context|$0.27/M input|$0.76/M output
Gemma
gemma 4 E4B it
google/gemma-4-e4b-it
Private

gemma 4 E4B it served on DeepInfra serverless inference.

Private|131K context|$0.02/M input|$0.10/M output
DeepInfra
MythoMax L2 13b
gryphe/mythomax-l2-13b
Private

MythoMax L2 13b served on DeepInfra serverless inference.

Private|4K context|$0.40/M input|$0.40/M output
DeepInfra
granite 4.2 30b
ibm-granite/granite-4.2-30b
Private

Granite-4.2-30B is the flagship reasoning model in the Granite 4.2 family. It delivers the strongest performance across reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Private|131K context|$0.16/M input|$0.65/M output
DeepInfra
granite 4.2 3b
ibm-granite/granite-4.2-3b
Private

Granite-4.2-3B is the compact reasoning model in the Granite 4.2 family. Despite its small parameter count, it delivers strong performance on reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Private|131K context|$0.03/M input|$0.12/M output
DeepInfra
granite 4.2 8b
ibm-granite/granite-4.2-8b
Private

Granite-4.2-8B is the mid-size reasoning model in the Granite 4.2 family. It delivers strong performance on reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Private|131K context|$0.06/M input|$0.25/M output
DeepInfra
Ling 3.0 flash
inclusionai/ling-3.0-flash
Private

The model prioritizes token efficiency and agentic inference at production scale, stretching what developers can achieve within limited token, latency, and serving-cost budgets.

Private|131K context|$0.06/M input|$0.18/M output
DeepInfra
Ling 3.0 flash Fin
inclusionai/ling-3.0-flash-fin
Private

Ling-3.0-flash-Fin is the first finance-enhanced model in the Ant Ling family. Developed by Ant Group with leading financial institutions and domain experts, it extends Ling-3.0-flash through continued training on high-quality financial data.

Private|262K context|$0.06/M input|$0.18/M output
Meta
Llama 3.3 70B Instruct Turbo
meta-llama/llama-3.3-70b-instruct-turbo
Private

Llama 3.3-70B Turbo is a highly optimized version of the Llama 3.3-70B model, utilizing FP8 quantization to deliver significantly faster inference speeds with a minor trade-off in accuracy. The model is designed to be helpful, safe, and flexible, with a focus on responsible deployment and mitigating potential risks such as bias, toxicity, and misinformation. It achieves state-of-the-art performance on various benchmarks, including conversational tasks, language translation, and text generation.

Private|131K context|$0.10/M input|$0.32/M output
Meta
Llama 4 Maverick 17B 128E Instruct FP8
meta-llama/llama-4-maverick-17b-128e-instruct-fp8
Private

The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Maverick, a 17 billion parameter model with 128 experts

Private|1M context|$0.20/M input|$0.80/M output
Meta
Llama 4 Scout 17B 16E Instruct
meta-llama/llama-4-scout-17b-16e-instruct
Private

The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Scout, a 17 billion parameter model with 16 experts

Private|328K context|$0.10/M input|$0.30/M output
Meta
Llama Guard 4 12B
meta-llama/llama-guard-4-12b
Private

Llama Guard 4 is a natively multimodal safety classifier with 12 billion parameters trained jointly on text and multiple images. Llama Guard 4 is a dense architecture pruned from the Llama 4 Scout pre-trained model and fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM inputs (prompt classification) and in LLM responses (response classification). It itself acts as an LLM: it generates text in its output that indicates whether a given prompt or response is safe or unsafe, and if unsafe, it also lists the content categories violated.

Private|164K context|$0.18/M input|$0.18/M output
Meta
Meta Llama 3.1 70B Instruct Turbo
meta-llama/meta-llama-3.1-70b-instruct-turbo
Private

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8B, 70B and 405B sizes

Private|131K context|$0.40/M input|$0.40/M output
Meta
Meta Llama 3.1 8B Instruct Turbo
meta-llama/meta-llama-3.1-8b-instruct-turbo
Private

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8B, 70B and 405B sizes

Private|131K context|$0.02/M input|$0.04/M output
Meta
Muse Glimmer 30B
meta-models/muse-glimmer-30b
Private

Muse Glimmer is a 30B multimodal agentic model distilled from Muse Spark — reasoning, tool use, and failure recovery in a single model that runs locally on consumer hardware.

Private|131K context|$0.30/M input|$1.20/M output
DeepInfra
phi 4
microsoft/phi-4
Private

Phi-4 is a model built upon a blend of synthetic datasets, data from filtered public domain websites, and acquired academic books and Q&A datasets. The goal of this approach was to ensure that small capable models were trained with data focused on high quality and advanced reasoning.

Private|16K context|$0.07/M input|$0.14/M output
Minimax
MiniMax M2.7 Turbo
minimax/minimax-m2.7-turbo
Anonymous

Speed-optimized MiniMax-M2.7

Anonymous|197K context|$0.38/M input|$1.70/M output
Minimax
MiniMax M3
minimax/minimax-m3
Private

MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.

Private|524K context|$0.28/M input|$1.10/M output
Mistral
Mistral Nemo Instruct 2407
mistralai/mistral-nemo-instruct-2407
Private

12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.

Private|131K context|$0.019/M input|$0.03/M output
Mistral
Mistral Small 24B Instruct 2501
mistralai/mistral-small-24b-instruct-2501
Private

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed for efficient local deployment. The model achieves 81% accuracy on the MMLU benchmark and performs competitively with larger models like Llama 3.3 70B and Qwen 32B, while operating at three times the speed on equivalent hardware.

Private|33K context|$0.05/M input|$0.08/M output
Mistral
Mistral Small 3.2 24B Instruct 2506
mistralai/mistral-small-3.2-24b-instruct-2506
Private

Mistral-Small-3.2-24B-Instruct is a drop-in upgrade over the 3.1 release, with markedly better instruction following, roughly half the infinite-generation errors, and a more robust function-calling interface—while otherwise matching or slightly improving on all previous text and vision benchmarks.

Private|128K context|$0.075/M input|$0.20/M output
Meta
Hermes 3 Llama 3.1 70B
nousresearch/hermes-3-llama-3.1-70b
Private

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the board.

Private|131K context|$0.70/M input|$0.70/M output
Nvidia
Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3b
Private

NVIDIA Nemotron 3 Nano is an open small reasoning model optimized for fast, cost-efficient inference in agentic and production workloads. Built with a hybrid Mixture-of-Experts (MoE) and Mamba-Transformer architecture, it delivers strong multi-step reasoning, high token throughput, stable latency with predictable cost, and efficient deployment for agent-based systems. Designed for real-world AI systems where reasoning can generate significantly more tokens per prompt, Nemotron Nano reduces compute cost while maintaining strong reasoning quality.

Private|262K context|$0.05/M input|$0.20/M output
Nvidia
Nemotron Content Safety 3.5
nvidia/nemotron-content-safety-3.5
Private

Nemotron Content Safety 3.5 is a multimodal safety classifier developed by NVIDIA. A compact safety model that handles text, images, and custom policies. It outputs a safe/unsafe classification plus a reasoning trace, and can be used as an inference-time guardrail, as a judge for LLM safety testing and evaluation, or with the accompanying training dataset to post-train models for safer behavior.

Private|131K context|$0.20/M input|$0.20/M output
Nvidia
NVIDIA Nemotron 3 Super 120B A12B
nvidia/nvidia-nemotron-3-super-120b-a12b
Private

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

Private|262K context|$0.085/M input|$0.40/M output
Nvidia
NVIDIA Nemotron 3.5 Lightning
nvidia/nvidia-nemotron-3.5-lightning
Private

NVIDIA Nemotron 3.5 Lightning is NVIDIA's fastest open model for always-on agents and high-volume specialized tasks. It delivers a substantial leap in agentic capability over its predecessor Nemotron 3 Nano, with up to 4x higher throughput on a 1M-token context.

Private|262K context|$0.08/M input|$0.20/M output
OpenAI
gpt oss 120b Turbo
openai/gpt-oss-120b-turbo
Private

gpt oss 120b Turbo served on DeepInfra serverless inference.

Private|131K context|$0.15/M input|$0.60/M output
OpenAI
gpt oss 120b Ultra
openai/gpt-oss-120b-ultra
Anonymous

Ultra speed version of gpt-oss-120b

Anonymous|131K context|$0.20/M input|$0.95/M output
OpenAI
OpenAI GPT OSS 20B
openai/gpt-oss-20b
Private

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.

Private|131K context|$0.03/M input|$0.14/M output
Qwen
Qwen2.5 72B Instruct
qwen/qwen2.5-72b-instruct
Private

Qwen2.5 is a model pretrained on a large-scale dataset of up to 18 trillion tokens, offering significant improvements in knowledge, coding, mathematics, and instruction following compared to its predecessor Qwen2. The model also features enhanced capabilities in generating long texts, understanding structured data, and generating structured outputs, while supporting multilingual capabilities for over 29 languages.

Private|33K context|$0.36/M input|$0.40/M output
Qwen
Qwen3 14B
qwen/qwen3-14b
Private

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support.

Private|41K context|$0.12/M input|$0.24/M output
Qwen
Qwen3 30B A3B
qwen/qwen3-30b-a3b
Private

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support

Private|41K context|$0.12/M input|$0.50/M output
Qwen
Qwen3 32B
qwen/qwen3-32b
E2EE

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support

E2EE|41K context|$0.08/M input|$0.28/M output
Qwen
Qwen3 Max
qwen/qwen3-max
Anonymous

The latest flagship model in the Qwen family. State-of-the-art results across a comprehensive suite of benchmarks — including knowledge, reasoning, coding, instruction following, human preference alignment, agent tasks, and multilingual understanding.

Anonymous|256K context|$1.20/M input|$6.00/M output
Qwen
Qwen3 Max Thinking
qwen/qwen3-max-thinking
Anonymous

The latest flagship reasoning model in the Qwen3 family. Further enhanced by multiple innovations like adaptive tool-use and advanced test-time scaling techniques

Anonymous|256K context|$1.20/M input|$6.00/M output
Qwen
Qwen3 VL 30B A3B Instruct
qwen/qwen3-vl-30b-a3b
E2EE

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities.

E2EE|262K context|$0.15/M input|$0.60/M output
Qwen
Qwen3.5 122B A10B
qwen/qwen3.5-122b-a10b
Private

Qwen3.5-122B-A10B is a large Mixture-of-Experts model from Alibaba's Qwen3.5 series with 122B total parameters and 10B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling, and support for 201 languages. Excels at complex reasoning, coding, multimodal understanding, and agentic tasks with the efficiency of sparse activation.

Private|262K context|$0.29/M input|$2.40/M output
Qwen
Qwen3.5 27B
qwen/qwen3.5-27b
Private

Qwen3.5-27B is Alibaba's largest dense Qwen3.5 model, delivering near-frontier quality across reasoning, coding, and instruction following. It features a 262K token context window (extensible to 1M), thinking/reasoning mode, tool calling, multi-token prediction, and support for 201 languages. Best suited for production deployments and complex enterprise tasks requiring top-tier performance.

Private|262K context|$0.26/M input|$2.60/M output
Qwen
Qwen3.8 27B
qwen/qwen3.8-27b
Private

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled.

Private|262K context|$0.40/M input|$3.00/M output
DeepInfra
L3 8B Lunaris v1 Turbo
sao10k/l3-8b-lunaris-v1-turbo
Private

L3 8B Lunaris v1 Turbo served on DeepInfra serverless inference.

Private|8K context|$0.04/M input|$0.05/M output
DeepInfra
L3.1 70B Euryale v2.2
sao10k/l3.1-70b-euryale-v2.2
Private

Euryale 3.1 - 70B v2.2 is a model focused on creative roleplay from Sao10k

Private|131K context|$0.85/M input|$0.85/M output
DeepInfra
Step 3.7 Flash
stepfun-ai/step-3.7-flash
Private

Step 3.7 Flash is an open-source multimodal reasoning model by StepFun with 198B total parameters (11B active) using Mixture of Experts. It accepts text and image inputs and features a 256K context window, selectable reasoning effort, tool calling, and agentic capabilities for coding and search workflows, scoring 80.9% on GPQA Diamond and 56.3% on SWE-bench Pro.

Private|262K context|$0.20/M input|$1.15/M output
Hy3
tencent/hy3
Private

Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Following the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post-training with higher quality data. Today, we introduce Hy3, which outperforms similar-size models and rivals flagship open-source models with 2-5x parameters. It also shows significant gains in utility across various products and productivity tasks.

Private|262K context|$0.14/M input|$0.58/M output
DeepInfra
Inkling Small
thinking-machines/inkling-small
Private

Inkling-Small is a Mixture-of-Experts transformer with 276B total parameters, 12B active, trained on NVIDIA GB300 NVL72 systems. Like Inkling, it features native reasoning over audio and images, variable thinking effort

Private|524K context|$0.45/M input|$1.20/M output
DeepInfra
MiMo V2.5
xiaomi/mimo-v2.5
Private

MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. Built upon the MiMo-V2-Flash backbone and extended with dedicated vision and audio encoders, it delivers robust performance across multimodal perception, long-context reasoning, and agentic workflows.

Private|262K context|$0.40/M input|$2.00/M output
DeepInfra
MiMo V2.5 Pro
xiaomi/mimo-v2.5-pro
Private

MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in [MiMo-V2-Flash](https://github.com/XiaomiMiMo/MiMo-V2-Flash).

Private|1M context|$1.00/M input|$3.00/M output
Nomic Embed Text
nomic-ai/nomic-embed-text
TEE

Nomic Embed Text served in a Tinfoil verified confidential enclave.

TEE|8K context|$0.05/M input|Verify/M output
Qwen
Qwen3.6 35B A3B Uncensored
qwen/qwen3.6-35b-a3b-uncensored
Private

Qwen3.6 35B A3B Uncensored served through the Phala AI private gateway.

Private|131K context|$0.30/M input|$1.50/M output
ace
ACE-Step 1.5
ace-step/ace-step-1.5
Anonymous

Feature-rich song generation with optional lyrics and detailed musical controls.

Anonymous|Unknown context||from $0.03/track
Alibaba
Wan 2.7 Pro Edit
alibaba/wan-2.7-pro-edit
Anonymous

Wan 2.7 Pro Edit is an image editing and inpainting model. High-fidelity results with strong prompt adherence and fast turnaround.

Anonymous|Unknown context||$0.09/edit
Flux
Flux 2 Max
black-forest-labs/flux-2-max-edit
Anonymous

Flux 2 Max is an image editing and inpainting model. Photorealistic, high-detail output with excellent typography and layout.

Anonymous|Unknown context||$0.12/edit
ByteDance
Seedream V4.5
bytedance/seedream-v4.5-edit
Anonymous

Seedream V4.5 is an image editing and inpainting model. Vivid color, sharp detail, and coherent multi-subject scenes.

Anonymous|Unknown context||$0.05/edit
ByteDance
Seedream V5 Lite
bytedance/seedream-v5-lite-edit
Anonymous

Seedream V5 Lite is an image editing and inpainting model. Vivid color, sharp detail, and coherent multi-subject scenes.

Anonymous|Unknown context||$0.05/edit
ByteDance
Seedream V5 Pro
bytedance/seedream-v5-pro
Anonymous

Seedream V5 Pro is a text-to-image model. Vivid color, sharp detail, and coherent multi-subject scenes.

Anonymous|Unknown context||from $0.06/image
ByteDance
Seedream V5 Pro
bytedance/seedream-v5-pro-edit
Anonymous

Seedream V5 Pro is an image editing and inpainting model. Vivid color, sharp detail, and coherent multi-subject scenes.

Anonymous|Unknown context||from $0.06/edit
ElevenLabs
ElevenLabs Multilingual v2
elevenlabs/elevenlabs-multilingual-v2
Anonymous

Multilingual text-to-speech using ElevenLabs. Supports 29 languages with high-quality natural-sounding voices, configurable speed, and accent accuracy.

Anonymous|Unknown context||$0.12/1K chars
ElevenLabs
ElevenLabs Music
elevenlabs/elevenlabs-music
Anonymous

High-quality instrumental music generation with configurable duration. Best for polished, production-ready tracks across a wide range of genres.

Anonymous|Unknown context||from $0.69/track
ElevenLabs
ElevenLabs Scribe V2
elevenlabs/elevenlabs-scribe-v2
Anonymous

ElevenLabs Scribe V2 is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.

Anonymous|Unknown context|$0.01/min|
ElevenLabs
ElevenLabs Sound Effects
elevenlabs/elevenlabs-sound-effects
Anonymous

Generate high-quality sound effects from text descriptions using ElevenLabs. Ideal for films, games, and digital content with configurable duration.

Anonymous|Unknown context||$0.14/min
ElevenLabs
ElevenLabs TTS v3
elevenlabs/elevenlabs-tts-v3
Anonymous

Generate natural text-to-speech audio using ElevenLabs Eleven-v3. High-quality voices with stability control and automatic text normalization.

Anonymous|Unknown context||$0.12/1K chars
DeepMind
Lyria 3 Pro
google/lyria-3-pro
Anonymous

Google's Lyria 3 Pro generates full-length, structured songs up to 3 minutes long from a single text prompt. Supports vocals, lyrics, and multi-language generation across genres.

Anonymous|Unknown context||$0.10/track
NanoBanana
Nano Banana 2
google/nano-banana-2
Anonymous

Nano Banana 2 is a text-to-image model. Creative, high-quality results with strong instruction following and clean edits.

Anonymous|Unknown context||from $0.10/image
NanoBanana
Nano Banana 2
google/nano-banana-2-edit
Anonymous

Nano Banana 2 is an image editing and inpainting model. Creative, high-quality results with strong instruction following and clean edits.

Anonymous|Unknown context||from $0.10/edit
NanoBanana
Nano Banana 2 Lite
google/nano-banana-2-lite-edit
Anonymous

Nano Banana 2 Lite is an image editing and inpainting model. Creative, high-quality results with strong instruction following and clean edits.

Anonymous|Unknown context||$0.06/edit
NanoBanana
Nano Banana Pro
google/nano-banana-pro
Anonymous

Nano Banana Pro is a text-to-image model. Creative, high-quality results with strong instruction following and clean edits.

Anonymous|Unknown context||from $0.18/image
NanoBanana
Nano Banana Pro
google/nano-banana-pro-edit
Anonymous

Nano Banana Pro is an image editing and inpainting model. Creative, high-quality results with strong instruction following and clean edits.

Anonymous|Unknown context||from $0.18/edit
Krea
Krea 2 Turbo
krea/krea-2-turbo
Private

Krea 2 Turbo is a text-to-image model. Aesthetic-focused, with a distinctive and stylized look.

Private|Unknown context||from $0.04/image
Luma
Luma Uni-1
luma/luma-uni-1-edit
Anonymous

Luma Uni-1 is an image editing and inpainting model. Cinematic lighting and composition with photoreal detail.

Anonymous|Unknown context||$0.06/edit
Luma
Luma Uni-1 Max
luma/luma-uni-1-max-edit
Anonymous

Luma Uni-1 Max is an image editing and inpainting model. Cinematic lighting and composition with photoreal detail.

Anonymous|Unknown context||$0.13/edit
Minimax
MiniMax Music 2.0
minimax/minimax-music-2.0
Anonymous

Full song generation with vocals and lyrics. Provide your own lyrics with verse/chorus structure for complete songs with singing.

Anonymous|Unknown context||$0.04/track
Minimax
MiniMax Music 2.5
minimax/minimax-music-2.5
Anonymous

Advanced song generation with vocals, lyrics optimizer, and instrumental mode. Supports structure tags and up to 3500 character lyrics.

Anonymous|Unknown context||$0.18/track
Minimax
MiniMax Music 2.6
minimax/minimax-music-2.6
Anonymous

Latest MiniMax song generation with vocals, instrumental mode, and support for rich structure tags in lyrics.

Anonymous|Unknown context||$0.18/track
Nvidia
Parakeet ASR
nvidia/parakeet-asr
Private

Parakeet ASR is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.

Private|Unknown context|$0.006/min|
OpenAI
GPT-5.6 Luna
openai/gpt-5.6-luna
Anonymous

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

Anonymous|1M context|$0.25/M input|$1.50/M output
OpenAI
GPT-5.6 Luna Pro
openai/gpt-5.6-luna-pro
Anonymous

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

Anonymous|1M context|$0.25/M input|$1.50/M output
OpenAI
GPT-5.6 Sol
openai/gpt-5.6-sol
Anonymous

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.

Anonymous|1M context|$2.50/M input|$12.50/M output
OpenAI
GPT Image 1.5
openai/gpt-image-1.5-edit
Anonymous

GPT Image 1.5 is an image editing and inpainting model. Instruction-following generation with broad world knowledge and reliable text rendering.

Anonymous|Unknown context||$0.31/edit
OpenAI
GPT Image 2
openai/gpt-image-2
Anonymous

GPT Image 2 is a text-to-image model. Instruction-following generation with broad world knowledge and reliable text rendering.

Anonymous|Unknown context||from $0.27/image
OpenAI
GPT Image 2
openai/gpt-image-2-edit
Anonymous

GPT Image 2 is an image editing and inpainting model. Instruction-following generation with broad world knowledge and reliable text rendering.

Anonymous|Unknown context||from $0.34/edit
OpenAI
Whisper Large V3
openai/whisper-large-v3
Private

Whisper Large V3 is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.

Private|Unknown context|$0.006/min|
OpenAI
Wizper (Whisper v3)
openai/wizper-whisper-v3
Private

Wizper (Whisper v3) is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.

Private|Unknown context|$0.006/min|
Qwen
Qwen 3.6 35B A3B FP8
qwen/qwen-3.6-35b-a3b-fp8
E2EE

Qwen 3.6 35B A3B FP8 running in a Trusted Execution Environment (TEE). A fast mixture-of-experts model with ~3B active parameters per token. Hardware attestation evidence is available for independent verification of enclave identity and configuration.

E2EE|32K context|$0.182/M input|$1.18/M output
Qwen
Qwen Edit Uncensored
qwen/qwen-edit-uncensored
Private

Qwen Edit Uncensored is an image editing and inpainting model. Flexible across a wide range of prompts and edits.

Private|Unknown context||$0.04/edit
Qwen
Qwen Image 2
qwen/qwen-image-2-edit
Anonymous

Qwen Image 2 is an image editing and inpainting model. Flexible across a wide range of prompts and edits.

Anonymous|Unknown context||$0.05/edit
Qwen
Qwen Image 2 Pro
qwen/qwen-image-2-pro-edit
Anonymous

Qwen Image 2 Pro is an image editing and inpainting model. Flexible across a wide range of prompts and edits.

Anonymous|Unknown context||$0.10/edit
Qwen
Qwen Image 3 Edit
qwen/qwen-image-3-edit
Anonymous

Qwen Image 3 Edit is an image editing and inpainting model. Flexible across a wide range of prompts and edits.

Anonymous|Unknown context||$0.04/edit
Qwen
Qwen Image 3 Pro
qwen/qwen-image-3-pro
Anonymous

Qwen Image 3 Pro is a text-to-image model. Flexible across a wide range of prompts and edits.

Anonymous|Unknown context||from $0.05/image
Qwen
Qwen Image 3 Pro Edit
qwen/qwen-image-3-pro-edit
Anonymous

Qwen Image 3 Pro Edit is an image editing and inpainting model. Flexible across a wide range of prompts and edits.

Anonymous|Unknown context||from $0.05/edit
Stability
Stable Audio 2.5
stabilityai/stable-audio-2.5
Anonymous

Fast, lightweight audio generation for sound effects, ambient textures, and short musical clips. Flexible duration from 5 seconds to over 3 minutes.

Anonymous|Unknown context||$0.19/track
Venice
FireRed Edit
venice/firered-edit
Private

FireRed Edit is an image editing and inpainting model.

Private|Unknown context||$0.04/edit
Venice
MMAudio V2
venice/mmaudio-v2
Anonymous

Generate synchronized audio and sound effects from text prompts with MMAudio V2.

Anonymous|Unknown context||$0.06/min
Venice
Muse Image
venice/muse-image-edit
Anonymous

Muse Image is an image editing and inpainting model.

Anonymous|Unknown context||$0.02/edit
Venice
Seed Audio 1.0
venice/seed-audio-1.0
Anonymous

Generate expressive multilingual speech and audio from a text prompt with BytePlus Seed Audio 1.0 (20 languages, timestamp length control).

Anonymous|Unknown context||$0.17/min
Venice
Sonilo V1.1 Music
venice/sonilo-v1.1-music
Anonymous

Generate licensed, commercial-use-safe music with precise control over style, mood, instrumentation, and duration.

Anonymous|Unknown context||$0.17/min
Venice
Sonilo V1.1 Sound Effects
venice/sonilo-v1.1-sound-effects
Anonymous

Generate licensed, commercial-use-safe sound effects with precise control over type, texture, intensity, and duration.

Anonymous|Unknown context||$0.12/min
Venice
Upscaler
venice/upscaler
Private

Upscaler is an image upscaling model.

Private|Unknown context||$0.01/upscale
Grok
Grok Imagine
x-ai/grok-imagine
Private

Grok Imagine is a text-to-image model.

Private|Unknown context||from $0.03/image
Grok
Grok Imagine 2.0
x-ai/grok-imagine-2.0
Private

Grok Imagine 2.0 is a text-to-image model.

Private|Unknown context||from $0.07/image
Grok
Grok Imagine 2.0
x-ai/grok-imagine-2.0-edit
Private

Grok Imagine 2.0 is an image editing and inpainting model.

Private|Unknown context||from $0.07/edit
Grok
Grok Imagine
x-ai/grok-imagine-edit
Private

Grok Imagine is an image editing and inpainting model.

Private|Unknown context||from $0.03/edit
Grok
Grok Imagine High Quality
x-ai/grok-imagine-high-quality
Private

Grok Imagine High Quality is an image editing and inpainting model.

Private|Unknown context||from $0.06/edit
Grok
Grok Imagine High Quality (SOTA)
x-ai/grok-imagine-high-quality-sota
Private

Grok Imagine High Quality (SOTA) is a text-to-image model.

Private|Unknown context||from $0.06/image
Grok
xAI Speech to Text v1
x-ai/xai-speech-to-text-v1
Anonymous

xAI Speech to Text v1 is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.

Anonymous|Unknown context|$0.00186/min|
Zhipu
GLM 4.7 Flash
z-ai/glm-4.7-flash
Private

GLM-4.7-Flash is a fast inference variant of GLM-4.7, optimized for speed while maintaining strong reasoning capabilities. Ideal for applications requiring quick responses with good quality.

Private|128K context|$0.06/M input|$0.40/M output
Mistral
Mistral Large 3 675B Instruct
mistralai/mistral-large-3-675b-instruct
Private

Mistral Large 3 675B Instruct through Amazon Bedrock Mantle with zero retention.

Private|256K context|$0.50/M input|$1.50/M output
Qwen
Qwen3 Coder Next
qwen/qwen3-coder-next
Private

Qwen3 Coder Next through Amazon Bedrock Mantle with zero retention.

Private|256K context|$0.50/M input|$1.20/M output

Privacy labels are metadata, not payload records

Pricing sources are shown per model. Privacy and moderation labels come from published provider metadata and stay separate from AnonRouter's own no-payload-log policy.

Review privacy
anonrouter - The private AI router