Bedrock

AWS Bedrock

inference provider · 47 models

Access 47 models served through AWS Bedrock on AnonRouter's privacy-first gateway, including Claude Opus 4.7, Claude Opus 4.8, and Claude Opus 5. Private Bedrock routes require zero-data-retention mode; anonymous Bedrock routes hide the end user's identity behind AnonRouter's AWS account but may retain payload content.

Models

47

Modalities

1

Text

From (input)

$0.04

per 1M tokens

Max context

1M

Private routes

47 / 47

not anonymous-only

Catalog by modality

47 routes
Text47

AWS Bedrock models47

Claude
Claude Opus 4.7
anthropic/claude-opus-4.7
Private

Claude Opus 4.7 is Anthropic's most capable generally available model for complex reasoning and agentic coding. It features a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.

Private|1M context|$5.50/M input|$27.50/M output
Claude
Claude Opus 4.8
anthropic/claude-opus-4.8
Private

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports long-horizon agentic work, complex multi-step coding, and memory-driven tasks where coherence over extended sessions matters. It features a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.

Private|1M context|$5.50/M input|$27.50/M output
Claude
Claude Opus 5
anthropic/claude-opus-5
Private

Claude Opus 5 is Anthropic's most capable model in the Opus family. It delivers major gains over Opus 4.8 in agentic coding, professional knowledge work, and long-horizon reasoning, with a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.

Private|1M context|$6.00/M input|$30.00/M output
Claude
Claude Sonnet 5
anthropic/claude-sonnet-5
Private

Claude Sonnet 5 is Anthropic's latest Sonnet model, substantially improving on Sonnet 4.6 in coding and agentic work and reaching near-Opus quality on many tasks. It features a 1M token context window, adaptive thinking, and strong document and vision understanding.

Private|1M context|$2.20/M input|$11.00/M output
DeepSeek
DeepSeek V3.2
deepseek/deepseek-v3.2
Private

DeepSeek-V3.2 is an efficient large language model with DeepSeek Sparse Attention (DSA) for long contexts. It features strong reasoning and tool-use skills, achieving top results on the 2025 IMO and IOI.

Private|164K context|$0.62/M input|$1.85/M output
Gemma
Google Gemma 3 27B Instruct
google/google-gemma-3-27b-instruct
Private

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to Gemma 2.

Private|128K context|$0.23/M input|$0.38/M output
Gemma
Google Gemma 4 26B A4B Instruct
google/google-gemma-4-26b-a4b-instruct
Private

Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.

Private|256K context|$0.13/M input|$0.40/M output
Gemma
Google Gemma 4 31B Instruct
google/google-gemma-4-31b-instruct
Private

Gemma 4 31B is a dense model from Google DeepMind with 31B parameters, delivering frontier-level reasoning performance. It handles text, image, and video input, supports 256K context, function calling, and configurable thinking modes.

Private|256K context|$0.14/M input|$0.40/M output
Minimax
MiniMax M2.5
minimax/minimax-m2.5
Private

MiniMax-M2.5 is a state-of-the-art large language model optimized for coding, agentic workflows, and modern application development with enhanced reasoning capabilities.

Private|196K context|$0.30/M input|$1.20/M output
Kimi
Kimi K2.5
moonshotai/kimi-k2.5
Private

Kimi K2.5 is Moonshot AIs most advanced open reasoning model, featuring trillion-parameter Mixture-of-Experts architecture with 32B active parameters and 256K context windows.

Private|256K context|$0.60/M input|$3.00/M output
Nvidia
NVIDIA Nemotron 3 Nano 30B
nvidia/nvidia-nemotron-3-nano-30b
Private

NVIDIA Nemotron 3 Nano 30B is a compact and efficient language model from NVIDIA, optimized for fast inference while maintaining strong performance across diverse tasks.

Private|256K context|$0.06/M input|$0.24/M output
OpenAI
GPT OSS 20B
openai/gpt-oss-20b
Private

GPT OSS 20B running in a Trusted Execution Environment (TEE). OpenAI's compact open-weight 21B MoE model with 3.6B active parameters, optimized for lower-latency inference, with hardware attestation evidence available for independent verification.

Private|128K context|$0.07/M input|$0.30/M output
OpenAI
OpenAI GPT OSS 120B
openai/openai-gpt-oss-120b
Private

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation

Private|128K context|$0.15/M input|$0.60/M output
Qwen
Qwen 3 235B A22B Instruct 2507
qwen/qwen-3-235b-a22b-instruct-2507
Private

Built for in-depth research and handling long, complex documents. Ideal for technical work, multimodal input, and high-precision tasks.

Private|256K context|$0.22/M input|$0.88/M output
Qwen
Qwen 3 Coder 480B Turbo
qwen/qwen-3-coder-480b-turbo
Private

Turbo variant of Qwen3 Coder 480B, optimized for faster inference on code tasks.

Private|128K context|$0.45/M input|$1.80/M output
Qwen
Qwen 3 Next 80b
qwen/qwen-3-next-80b
Private

Optimized for speed and efficiency.

Private|256K context|$0.14/M input|$1.20/M output
Qwen
Qwen3 VL 235B
qwen/qwen3-vl-235b
Private

Qwen3-VL 235B vision-language model with MoE architecture. The most powerful VL model in the Qwen series with superior visual perception, OCR, and multimodal reasoning.

Private|256K context|$0.53/M input|$2.66/M output
Grok
Grok 4.3
x-ai/grok-4.3
Private

Grok 4.3 is xAI's most intelligent and fastest reasoning model with function calling, structured outputs, and a 1M-token context window. Suited for agentic workflows, instruction-following tasks, and applications requiring high factual accuracy.

Private|1M context|$1.25/M input|$2.50/M output
Zhipu
GLM 4.6
z-ai/glm-4.6
Private

GLM-4.6 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.

Private|198K context|$0.43/M input|$1.75/M output
Zhipu
GLM 4.7
z-ai/glm-4.7
Private

GLM-4.7 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.

Private|203K context|$0.60/M input|$0.40/M output
Zhipu
GLM 4.7 Flash
z-ai/glm-4.7-flash
Private

GLM-4.7-Flash is a fast inference variant of GLM-4.7, optimized for speed while maintaining strong reasoning capabilities. Ideal for applications requiring quick responses with good quality.

Private|203K context|$0.07/M input|$0.40/M output
Zhipu
GLM 5
z-ai/glm-5
Private

GLM-5 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple languages. Supports large context windows for processing extensive text and detailed analysis.

Private|200K context|$1.00/M input|$3.20/M output
Gemma
Gemma 3 4B IT
google/gemma-3-4b-it
Private

Gemma 3 4B IT through Amazon Bedrock Mantle with zero retention.

Private|128K context|$0.04/M input|$0.08/M output
Mistral
Mistral Large 3 675B Instruct
mistralai/mistral-large-3-675b-instruct
Private

Mistral Large 3 675B Instruct through Amazon Bedrock Mantle with zero retention.

Private|256K context|$0.50/M input|$1.50/M output
Qwen
Qwen3 Coder Next
qwen/qwen3-coder-next
Private

Qwen3 Coder Next through Amazon Bedrock Mantle with zero retention.

Private|256K context|$0.50/M input|$1.20/M output
Claude
Claude Haiku 4.5
anthropic/claude-haiku-4-5
Private

Claude Haiku 4.5 served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|200K context|$1.10/M input|$5.50/M output
DeepSeek
DeepSeek V3.1
deepseek/deepseek-v3-1
Private

DeepSeek V3.1 served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.58/M input|$1.68/M output
Gemma
Gemma 3 12B IT
google/gemma-3-12b-it
Private

Gemma 3 12B IT served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.09/M input|$0.29/M output
Gemma
Gemma 4 E2B
google/gemma-4-e2b
Private

Gemma 4 E2B served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.04/M input|$0.08/M output
Minimax
MiniMax M2
minimax/minimax-m2
Private

MiniMax M2 served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|1M context|$0.30/M input|$1.20/M output
Minimax
MiniMax M2.1
minimax/minimax-m2-1
Private

MiniMax M2.1 served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|196K context|$0.30/M input|$1.20/M output
Mistral
Devstral 2 123B
mistralai/devstral-2-123b
Private

Devstral 2 123B served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|256K context|$0.40/M input|$2.00/M output
Mistral
Magistral Small 2509
mistralai/magistral-small-2509
Private

Magistral Small 2509 served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.50/M input|$1.50/M output
Mistral
Ministral 3 3B Instruct
mistralai/ministral-3-3b-instruct
Private

Ministral 3 3B Instruct served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.10/M input|$0.10/M output
Mistral
Ministral 3 8B Instruct
mistralai/ministral-3-8b-instruct
Private

Ministral 3 8B Instruct served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.15/M input|$0.15/M output
Mistral
Ministral 3 14B Instruct
mistralai/ministral-3-14b-instruct
Private

Ministral 3 14B Instruct served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.20/M input|$0.20/M output
Mistral
Voxtral Mini 3B 2507
mistralai/voxtral-mini-3b-2507
Private

Voxtral Mini 3B 2507 served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|32K context|$0.04/M input|$0.04/M output
Mistral
Voxtral Small 24B 2507
mistralai/voxtral-small-24b-2507
Private

Voxtral Small 24B 2507 served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|32K context|$0.10/M input|$0.30/M output
Kimi
Kimi K2 Thinking
moonshotai/kimi-k2-thinking
Private

Kimi K2 Thinking served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|256K context|$0.60/M input|$2.50/M output
Nvidia
Nemotron Nano 9B V2
nvidia/nemotron-nano-9b-v2
Private

Nemotron Nano 9B V2 served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.06/M input|$0.23/M output
Nvidia
Nemotron Nano 12B V2
nvidia/nemotron-nano-12b-v2
Private

Nemotron Nano 12B V2 served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.20/M input|$0.60/M output
Nvidia
Nemotron Super 3 120B
nvidia/nemotron-super-3-120b
Private

Nemotron Super 3 120B served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|256K context|$0.15/M input|$0.65/M output
OpenAI
GPT-OSS Safeguard 20B
openai/gpt-oss-safeguard-20b
Private

GPT-OSS Safeguard 20B served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.07/M input|$0.20/M output
OpenAI
GPT-OSS Safeguard 120B
openai/gpt-oss-safeguard-120b
Private

GPT-OSS Safeguard 120B served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|128K context|$0.15/M input|$0.60/M output
Qwen
Qwen3 32B
qwen/qwen3-32b
Private

Qwen3 32B served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|32K context|$0.15/M input|$0.60/M output
Qwen
Qwen3 Coder 30B A3B Instruct
qwen/qwen3-coder-30b-a3b-instruct
Private

Qwen3 Coder 30B A3B Instruct served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|256K context|$0.15/M input|$0.60/M output
Bedrock
Palmyra Vision 7B
writer/palmyra-vision-7b
Private

Palmyra Vision 7B served through Amazon Bedrock with support for AWS zero-data-retention mode.

Private|4K context|$0.15/M input|$0.60/M output
Provider documentation44 sources
AWS Bedrock Models APIAvailability verified in us-east-1 on 2026-07-31; Private entries require allowed_modes to include none, while Anonymous entries may permit payload retention.AWS Bedrock data protectionAWS says model providers do not have access to Bedrock logs, customer prompts, or completions.AWS Bedrock data retentionAWS documents configurable data retention modes, including zero data retention where supported.AWS invocation loggingAWS model invocation logging can collect inputs/outputs when enabled and is disabled by default.AWS Bedrock FAQAWS FAQ states customer Bedrock inputs and outputs are not used to train Amazon Nova, Amazon Titan, or third-party models.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS Bedrock pricingOfficial standard on-demand token pricing.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS Price List APIStandard-tier us-east-1 prices were read from the authenticated AWS Price List API.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.AWS model cardOfficial context window, maximum output, launch date, and modality source.
AWS Bedrock provider — AnonRouter