Cheapest AI Models

collections/low-cost · 49 models

Text routes priced at or below $0.15 per million input tokens. Cheap is not the opposite of private here: several of these run inside enclaves, because the open-weight models that are inexpensive to serve are also the ones providers can host confidentially.

Models

49

Labs

15

model creators

From (input)

$0.019

per 1M tokens

Max context

1M

Private routes

45 / 49

up to E2EE

In this collection49

Ordered by lowest input price.

Mistral
Mistral Nemo Instruct 2407
mistralai/mistral-nemo-instruct-2407
Private

12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.

Private|131K context|$0.019/M input|$0.03/M output
Gemma
gemma 4 E4B it
google/gemma-4-e4b-it
Private

gemma 4 E4B it served on DeepInfra serverless inference.

Private|131K context|$0.02/M input|$0.10/M output
Meta
Meta Llama 3.1 8B Instruct Turbo
meta-llama/meta-llama-3.1-8b-instruct-turbo
Private

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8B, 70B and 405B sizes

Private|131K context|$0.02/M input|$0.04/M output
DeepInfra
granite 4.2 3b
ibm-granite/granite-4.2-3b
Private

Granite-4.2-3B is the compact reasoning model in the Granite 4.2 family. Despite its small parameter count, it delivers strong performance on reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Private|131K context|$0.03/M input|$0.12/M output
OpenAI
OpenAI GPT OSS 20B
openai/gpt-oss-20b
Private

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.

Private|131K context|$0.03/M input|$0.14/M output
DeepInfra
L3 8B Lunaris v1 Turbo
sao10k/l3-8b-lunaris-v1-turbo
Private

L3 8B Lunaris v1 Turbo served on DeepInfra serverless inference.

Private|8K context|$0.04/M input|$0.05/M output
Gemma
gemma 3 12b it
google/gemma-3-12b-it
Private

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3-12B is Google's latest open source model, successor to Gemma 2

Private|131K context|$0.05/M input|$0.15/M output
Gemma
gemma 3 4b it
google/gemma-3-4b-it
Private

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3-12B is Google's latest open source model, successor to Gemma 2

Private|131K context|$0.05/M input|$0.10/M output
Inception
Mercury 2.5
inception/mercury-2.5
Anonymous

Mercury 2.5 is a diffusion-based reasoning model from Inception with fast parallel token generation, tunable reasoning, tool calling, and structured output support.

Anonymous|260K context|$0.05/M input|$0.1875/M output
Mistral
Mistral Small 24B Instruct 2501
mistralai/mistral-small-24b-instruct-2501
Private

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed for efficient local deployment. The model achieves 81% accuracy on the MMLU benchmark and performs competitively with larger models like Llama 3.3 70B and Qwen 32B, while operating at three times the speed on equivalent hardware.

Private|33K context|$0.05/M input|$0.08/M output
Nvidia
Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3b
Private

NVIDIA Nemotron 3 Nano is an open small reasoning model optimized for fast, cost-efficient inference in agentic and production workloads. Built with a hybrid Mixture-of-Experts (MoE) and Mamba-Transformer architecture, it delivers strong multi-step reasoning, high token throughput, stable latency with predictable cost, and efficient deployment for agent-based systems. Designed for real-world AI systems where reasoning can generate significantly more tokens per prompt, Nemotron Nano reduces compute cost while maintaining strong reasoning quality.

Private|262K context|$0.05/M input|$0.20/M output
Qwen
Qwen 2.5 7B
qwen/qwen-2.5-7b
E2EE

Qwen 2.5 7B Instruct running in a Trusted Execution Environment (TEE). A compact model with strong coding, math, and multilingual capabilities supporting 29+ languages, with hardware attestation evidence available for independent verification.

E2EE|32K context|$0.05/M input|$0.13/M output
DeepSeek
DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
Private

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.

Private|1M context|$0.06/M input|$0.18/M output
DeepInfra
granite 4.2 8b
ibm-granite/granite-4.2-8b
Private

Granite-4.2-8B is the mid-size reasoning model in the Granite 4.2 family. It delivers strong performance on reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Private|131K context|$0.06/M input|$0.25/M output
DeepInfra
Ling 3.0 flash
inclusionai/ling-3.0-flash
Private

The model prioritizes token efficiency and agentic inference at production scale, stretching what developers can achieve within limited token, latency, and serving-cost budgets.

Private|131K context|$0.06/M input|$0.18/M output
DeepInfra
Ling 3.0 flash Fin
inclusionai/ling-3.0-flash-fin
Private

Ling-3.0-flash-Fin is the first finance-enhanced model in the Ant Ling family. Developed by Ant Group with leading financial institutions and domain experts, it extends Ling-3.0-flash through continued training on high-quality financial data.

Private|262K context|$0.06/M input|$0.18/M output
DeepInfra
Ling 3.0 flash VL
inclusionai/ling-3.0-flash-vl
Private

The multimodal version built on Ling-3.0-flash — 124B total / ~5.5B active per token, with native text, image, and video understanding. It’s mainly designed for multimodal agentic workflows, long-context understanding, and multi-step reasoning.

Private|131K context|$0.06/M input|$0.18/M output
Zhipu
GLM 4.7 Flash
z-ai/glm-4.7-flash
Private

GLM-4.7-Flash is a fast inference variant of GLM-4.7, optimized for speed while maintaining strong reasoning capabilities. Ideal for applications requiring quick responses with good quality.

Private|128K context|$0.06/M input|$0.40/M output
Gemma
gemma 4 26B A4B it
google/gemma-4-26b-a4b-it
Private

Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

Private|262K context|$0.07/M input|$0.34/M output
DeepInfra
phi 4
microsoft/phi-4
Private

Phi-4 is a model built upon a blend of synthetic datasets, data from filtered public domain websites, and acquired academic books and Q&A datasets. The goal of this approach was to ensure that small capable models were trained with data focused on high quality and advanced reasoning.

Private|16K context|$0.07/M input|$0.14/M output
Zhipu
GLM 4.7 Flash Heretic
z-ai/glm-4.7-flash-heretic
Private

GLM-4.7-Flash-Heretic is an uncensored experimental variant of GLM-4.7-Flash, optimized for creative freedom and unfiltered dialogue with fast inference speed.

Private|200K context|$0.07/M input|$0.40/M output
Mistral
Mistral Small 3.2 24B Instruct 2506
mistralai/mistral-small-3.2-24b-instruct-2506
Private

Mistral-Small-3.2-24B-Instruct is a drop-in upgrade over the 3.1 release, with markedly better instruction following, roughly half the infinite-generation errors, and a more robust function-calling interface—while otherwise matching or slightly improving on all previous text and vision benchmarks.

Private|128K context|$0.075/M input|$0.20/M output
Nvidia
NVIDIA Nemotron 3 Nano 30B
nvidia/nvidia-nemotron-3-nano-30b
Private

NVIDIA Nemotron 3 Nano 30B is a compact and efficient language model from NVIDIA, optimized for fast inference while maintaining strong performance across diverse tasks.

Private|128K context|$0.075/M input|$0.30/M output
Gemma
gemma 3 27b it
google/gemma-3-27b-it
Private

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to Gemma 2

Private|131K context|$0.08/M input|$0.16/M output
Nvidia
NVIDIA Nemotron 3.5 Lightning
nvidia/nvidia-nemotron-3.5-lightning
Private

NVIDIA Nemotron 3.5 Lightning is NVIDIA's fastest open model for always-on agents and high-volume specialized tasks. It delivers a substantial leap in agentic capability over its predecessor Nemotron 3 Nano, with up to 4x higher throughput on a 1M-token context.

Private|262K context|$0.08/M input|$0.20/M output
Qwen
Qwen3 32B
qwen/qwen3-32b
Private

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support

Private|41K context|$0.08/M input|$0.28/M output
Nvidia
NVIDIA Nemotron 3 Super 120B A12B
nvidia/nvidia-nemotron-3-super-120b-a12b
Private

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

Private|262K context|$0.085/M input|$0.40/M output
DeepSeek
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
Private

DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks.

Private|1M context|$0.09/M input|$0.18/M output
Gemma
gemma 4 31B it turbo
google/gemma-4-31b-it-turbo
Private

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

Private|262K context|$0.09/M input|$0.34/M output
Mistral
Mistral Small 3.2 24B Instruct
mistralai/mistral-small-3.2-24b-instruct
Private

Mistral Small 3.2 is a 24B parameter model optimized for efficiency and performance. Ideal for general-purpose tasks with balanced speed and capability.

Private|256K context|$0.0938/M input|$0.25/M output
DeepInfra
Seed 2.0 mini
bytedance/seed-2.0-mini
Anonymous

Built for low-latency, high-concurrency, cost-sensitive use cases, with flexible deployment, four-tier thinking, and multimodal

Anonymous|256K context|$0.10/M input|$0.40/M output
Meta
Llama 3.3 70B Instruct Turbo
meta-llama/llama-3.3-70b-instruct-turbo
Private

Llama 3.3-70B Turbo is a highly optimized version of the Llama 3.3-70B model, utilizing FP8 quantization to deliver significantly faster inference speeds with a minor trade-off in accuracy. The model is designed to be helpful, safe, and flexible, with a focus on responsible deployment and mitigating potential risks such as bias, toxicity, and misinformation. It achieves state-of-the-art performance on various benchmarks, including conversational tasks, language translation, and text generation.

Private|131K context|$0.10/M input|$0.32/M output
Meta
Llama 4 Scout 17B 16E Instruct
meta-llama/llama-4-scout-17b-16e-instruct
Private

The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Scout, a 17 billion parameter model with 16 experts

Private|328K context|$0.10/M input|$0.30/M output
Qwen
Qwen 3.5 9B
qwen/qwen-3.5-9b
Private

A 9B dense model with 262K native context window (extendable to 1M). Features Gated DeltaNet hybrid attention architecture for efficient long-context processing. Supports 201 languages, thinking/reasoning mode, and function calling.

Private|256K context|$0.10/M input|$0.15/M output
Qwen
Qwen 3.6 35B A3B
qwen/qwen-3.6-35b-a3b
Private

Qwen 3.6 35B A3B is a fast mixture-of-experts model with 35B total parameters and ~3B active per token. Strong at agentic coding, STEM reasoning, and tool use, with a native 256K context window.

Private|256K context|$0.10/M input|$1.00/M output
Qwen
Qwen3.8 Flash
qwen/qwen3.8-flash
Anonymous

Qwen's fast, low-cost model with a one-million-token context window, billed at a flat rate across the entire window.

Anonymous|1M context|$0.113/M input|$0.382/M output
Gemma
Google Gemma 3 27B Instruct
google/google-gemma-3-27b-instruct
Private

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to Gemma 2.

Private|198K context|$0.12/M input|$0.20/M output
Qwen
Qwen3 14B
qwen/qwen3-14b
Private

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support.

Private|41K context|$0.12/M input|$0.24/M output
Qwen
Qwen3 30B A3B
qwen/qwen3-30b-a3b
Private

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support

Private|41K context|$0.12/M input|$0.50/M output
Gemma
Google Gemma 4 26B A4B Instruct
google/google-gemma-4-26b-a4b-instruct
Private

Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.

Private|256K context|$0.13/M input|$0.40/M output
DeepSeek
DeepSeek V4 Flash 0423
deepseek/deepseek-v4-flash-0423
Private

DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.

Private|1M context|$0.138/M input|$0.275/M output
Qwen
Qwen 3.8 Flash
qwen/qwen-3.8-flash
Anonymous

Qwen 3.8 Flash is the latest multimodal model in the Qwen family, pairing strong reasoning and generation with remarkable speed. It natively supports a 1M-token context window, so it can process long documents, entire codebases, and complex conversations in a single pass. It excels at coding assistance, agentic workflows, and visual understanding, accepting text, image, and video input, and its thinking mode can be turned on or off per request.

Anonymous|1M context|$0.14/M input|$0.49/M output
Hy3
tencent/hy3
Private

Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Following the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post-training with higher quality data. Today, we introduce Hy3, which outperforms similar-size models and rivals flagship open-source models with 2-5x parameters. It also shows significant gains in utility across various products and productivity tasks.

Private|262K context|$0.14/M input|$0.58/M output
Meta
Llama 3.2 3B
meta-llama/llama-3.2-3b
Private

Llama 3.2 3B is a text model.

Private|128K context|$0.15/M input|$0.60/M output
OpenAI
OpenAI GPT OSS 120B
openai/gpt-oss-120b
TEE

OpenAI GPT OSS 120B served in a Tinfoil verified confidential enclave.

TEE|131K context|$0.15/M input|$0.60/M output
OpenAI
gpt oss 120b Turbo
openai/gpt-oss-120b-turbo
Private

gpt oss 120b Turbo served on DeepInfra serverless inference.

Private|131K context|$0.15/M input|$0.60/M output
Qwen
Qwen 3 235B A22B Instruct 2507
qwen/qwen-3-235b-a22b-instruct-2507
Private

Built for in-depth research and handling long, complex documents. Ideal for technical work, multimodal input, and high-precision tasks.

Private|128K context|$0.15/M input|$0.75/M output
Qwen
Qwen3 VL 30B A3B Instruct
qwen/qwen3-vl-30b-a3b
Private

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities.

Private|262K context|$0.15/M input|$0.60/M output
Zhipu
GLM 5.3 Flash
z-ai/glm-5.3-flash
Private

GLM-5.3 Flash is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.

Private|1M context|$0.15/M input|$0.50/M output

This list is rebuilt from the live catalog rather than stored as a snapshot, so it tracks pricing, context windows, and privacy tiers as providers change them. Ordering is yours to pick, and there is no popularity option: prompts are never retained, and the usage metadata kept for billing is not turned into a public ranking.

Explore more collections

Strongest guarantee in this collection:E2EE

Cheapest AI Models - AnonRouter