DeepSeek V3.2 API pricing
9 providers serve DeepSeek V3.2. DeepSeek V3.2 is an open-weight model (MIT) with 685B total parameters, 37B active per token (mixture-of-experts), up to 164K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-08-23 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| Novita AI | $0.269 | $0.400output floor | $0.135 | 164K | Use |
| DeepSeek | $0.280 | $0.400 | - | 164K | Use |
| GMI Cloud | $0.280 | $0.400 | - | 164K | Use |
| OpenRouter | $0.280 | $0.400 | - | 164K | Use |
| Fireworks AI | $0.560 | $1.68 | - | 164K | Use |
| Google Vertex AI | $0.560 | $1.68 | - | 164K | Use |
| Azure AI Foundry | $0.580 | $1.68 | - | 164K | Use |
| AWS Bedrock | $0.620 | $1.85 | - | 164K | Use |
| SambaNova | $3.00 | $4.50 | - | 33K | Use |
The cheapest way to run DeepSeek V3.2
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $30.83/month, at Novita AI. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, the API floor remains below our self-hosting estimate of ~$1.27/1M. Details below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: DeepSeek V3.2 self-hosting economics
DeepSeek V3.2 is open-weight (MIT), so the API price above competes with the GPU-hour market. 685B total parameters (MoE, ~37B active per token) needs roughly 942 GB VRAM at FP8 or 471 GB at INT4, KV-cache headroom included.
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| DeepSeek R1 | $0.200 | $0.400 | 17 |
| DeepSeek V3 | $0.200 | $0.200 | 16 |
| DeepSeek R1 Distill Llama 70B | $0.200 | $0.375 | 10 |
| DeepSeek-V3.1 | $0.270 | $1.00 | 9 |
| DeepSeek V4 Flash | $0.050 | $0.101 | 9 |
| DeepSeek V4 Pro | $0.435 | $0.870 | 6 |
FAQ
What is the cheapest API for DeepSeek V3.2?
As of 2026-08-23, the lowest input price for DeepSeek V3.2 is Novita AI at $0.269 per 1M input tokens. The lowest output price, $0.400 per 1M tokens, is likewise shared by Novita AI, DeepSeek, GMI Cloud and OpenRouter. The cheapest provider for a real workload depends on its input/output mix. The next-lowest input price is DeepSeek at $0.280, a 4% difference.
How much VRAM do you need to self-host DeepSeek V3.2?
DeepSeek V3.2 has 685B parameters, so plan for roughly 942 GB of VRAM at FP8 or 471 GB at INT4/AWQ, KV-cache headroom included. That exceeds a single GPU. A typical node is 5× MI300X.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON