DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.
Modalities
In / out price
$0.08 / $0.18
per 1M
Cached price
$0.016
per 1M
Context
1M
Max output
16K
Released
Jul 31, 2026
Providers
3 routesThe same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.
Privacy
2 / 4Private routing
Private tier
- Anonymous
- Private
- TEE
- E2EE
Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed. Learn more