Ling 3.0 flash VL
The multimodal version built on Ling-3.0-flash — 124B total / ~5.5B active per token, with native text, image, and video understanding. It’s mainly designed for multimodal agentic workflows, long-context understanding, and multi-step reasoning.
This DeepInfra route is not in the reviewed endpoint manifest and cannot be enabled.
Modalities
In / out price
$0.06 / $0.18
per 1M
Cached price
$0.012
per 1M
Context
131K
Max output
16K
Released
–
Providers
1 routeThe same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.
| Provider | Privacy | Input / 1M | Output / 1M | Cache read / 1M | Context | Max output | Uptime | Latency | Throughput |
|---|---|---|---|---|---|---|---|---|---|
DeepInfraNot live | Private | $0.06 | $0.18 | $0.012 | 131K | 16K | — | — | — |
Privacy
2 / 4Private routing
Private tier
- Anonymous
- Private
- TEE
- E2EE
Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed. Learn more