Mercury 2 is a diffusion-based reasoning LLM from Inception, delivering over 1,000 tokens per second — 5x faster than leading speed-optimized models — with strong reasoning, tool use, and structured output capabilities.
Modalities
In / out price
$0.3125 / $0.9375
per 1M
Cached price
$0.0313
per 1M
Context
128K
Max output
50K
Released
Feb 20, 2026
Providers
1 routeThe same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.
| Provider | Privacy | Input / 1M | Output / 1M | Cache read / 1M | Context | Max output | Uptime | Latency | Throughput |
|---|---|---|---|---|---|---|---|---|---|
| Anonymous | $0.3125 | $0.9375 | $0.0313 | 128K | 50K | — | — | — |
Privacy
1 / 4Anonymous routing
Anonymous tier
- Anonymous
- Private
- TEE
- E2EE
Your identity is hidden from the inference provider. The inference provider can see prompt content; zero retention is not guaranteed. Learn more
Features
6Streaming
Tool calling
Reasoning
JSON/schema
Web search
Prompt caching