DeepSeek V4 Flash API pricing
22 providers serve DeepSeek V4 Flash. DeepSeek V4 Flash is an open-weight model (MIT) with 284B total parameters, 13B active per token (mixture-of-experts), up to 1049K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-10-08 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| OpenRouter | $0.030 | $1.28 | $0.030 | 1049K | Use |
| Sail | $0.090 | $0.180output floor | $0.020 | 1049K | |
| DeepInfra | $0.090 | $0.180 | $0.018 | 1049K | Use |
| Pinstripes | $0.100 | $0.200 | - | 164K | |
| Baseten | $0.130 | $0.260 | $0.028 | 1049K | Use |
| Databricks | $0.140 | $0.280 | $0.028 | 1000K | Use |
| Fireworks AI | $0.140 | $0.280 | $0.028 | 1049K | Use |
| Nebius | $0.140 | $0.280 | - | 1049K | Use |
| Together AI | $0.140 | $0.280 | $0.030 | 1049K | Use |
| Tensormesh | $0.140 | $0.280 | - | 33K | |
| Tencent | $0.140 | $0.280 | $0.0028 | 1000K | |
| Novita AI | $0.140 | $0.280 | $0.028 | 1049K | Use |
| W&B Inference | $0.140 | $0.280 | $0.070 | 1049K | Use |
| Aihubmix | $0.142 | $0.284 | $0.028 | 1000K | |
| Prism | $0.170 | $0.210 | $0.070 | 1000K | |
| Azure AI Foundry | $0.190 | $0.510 | $0.028 | 1000K | Use |
| Alibaba Model Studio | $0.200 | $0.400 | $0.040 | 1000K | Use |
| Qwencloud | $0.200 | $0.400 | $0.040 | 1000K | |
| Qwen AI Platform | $0.200 | $0.400 | $0.040 | 1000K | |
| Libertai | $0.250 | $1.75 | - | 200K | |
| DeepSeek | $0.300 | $1.20 | $0.0060 | 1000K | Use |
| Scaleway | $0.400 | $0.800 | $0.080 | 256K | Use |
The cheapest way to run DeepSeek V4 Flash
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $11.70/month — the same at Sail and DeepInfra. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, the API floor remains below our self-hosting estimate of ~$0.209/1M. Details below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: DeepSeek V4 Flash self-hosting economics
DeepSeek V4 Flash is open-weight (MIT), so the API price above competes with the GPU-hour market. 284B total parameters (MoE, ~13B active per token) needs roughly 391 GB VRAM at FP8 or 196 GB at INT4, KV-cache headroom included.
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| DeepSeek V4 Pro | $0.209 | $0.418 | 17 |
| DeepSeek R1 | $0.200 | $0.400 | 17 |
| DeepSeek V3 | $0.200 | $0.200 | 15 |
| DeepSeek V3.2 | $0.260 | $0.380 | 11 |
| DeepSeek-V3.1 | $0.250 | $0.950 | 11 |
| DeepSeek R1 Distill Llama 70B | $0.200 | $0.375 | 10 |
FAQ
What is the cheapest API for DeepSeek V4 Flash?
As of 2026-10-08, the lowest input price for DeepSeek V4 Flash is OpenRouter at $0.030 per 1M input tokens. The lowest output price, $0.180 per 1M tokens, is likewise shared by Sail and DeepInfra. The cheapest provider for a real workload depends on its input/output mix. The next-lowest input price is Sail at $0.090, a 200% difference.
How much VRAM do you need to self-host DeepSeek V4 Flash?
DeepSeek V4 Flash has 284B parameters, so plan for roughly 391 GB of VRAM at FP8 or 196 GB at INT4/AWQ, KV-cache headroom included. That exceeds a single GPU. A typical node is 3× H200 NVL.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON