DeepSeek

DeepSeek V4 Flash 0731

deepseek
Chat

DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.

Modalities

In / out price

$0.08 / $0.18

per 1M

Cached price

$0.016

per 1M

Context

1M

Max output

16K

Released

Jul 31, 2026

Providers

3 routes

The same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.

ProviderPrivacyInput / 1MOutput / 1MCache read / 1MContextMax outputUptimeLatencyThroughput
VeniceVeniceNot live
Private$0.175$0.35$0.0351M33K
DeepInfraDeepInfraNot live
Private$0.08$0.18$0.0161M16K
Phala AINot live
Private$0.44$1.32$0.0281M393K

Privacy

2 / 4

Private routing

Private tier

  1. Anonymous
  2. Private
  3. TEE
  4. E2EE

Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed. Learn more

Features

7
Streaming
Tool calling
Reasoning
JSON/schema
Web search
Code optimized
Prompt caching

Uptime

DeepSeek V4 Flash 0731 — AnonRouter Models