DeepSeek R1 Distill Llama 70B API pricing
10 providers serve DeepSeek R1 Distill Llama 70B. DeepSeek R1 Distill Llama 70B is an open-weight model (MIT) with 70B total parameters, up to 131K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-08-23 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| DeepInfra | $0.200 | $0.600 | - | 131K | Use |
| Nebius | $0.250 | $0.750 | - | 128K | Use |
| Nscale | $0.375 | $0.375output floor | - | - | Use |
| OVHcloud | $0.670 | $0.670 | - | 131K | Use |
| SambaNova | $0.700 | $1.40 | - | 131K | Use |
| Vercel AI Gateway | $0.750 | $0.990 | - | 131K | Use |
| Novita AI | $0.800 | $0.800 | - | 8K | Use |
| OpenRouter marketplace quote | $0.800 | $0.800 | - | 8K | Use |
| Fireworks AI | $0.900 | $0.900 | - | 131K | Use |
| DigitalOcean (Paperspace) | $0.990 | $0.990 | - | 33K | Use |
The cheapest way to run DeepSeek R1 Distill Llama 70B
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $32.00/month, at DeepInfra. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, the API floor remains below our self-hosting estimate of ~$0.376/1M. Details below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: DeepSeek R1 Distill Llama 70B self-hosting economics
DeepSeek R1 Distill Llama 70B is open-weight (MIT), so the API price above competes with the GPU-hour market. 70B total parameters needs roughly 97 GB VRAM at FP8 or 49 GB at INT4, KV-cache headroom included.
| GPU that fits (FP8, single card) | VRAM | From $/hr | Cheapest at | |
|---|---|---|---|---|
| H200 NVL | 141 GB | $0.500community | RunPod | Rent |
| MI300X | 192 GB | $0.500community | RunPod | Rent |
| H200 | 141 GB | $2.00spot | Verda | |
| B200 | 180 GB | $3.06spot | Verda | |
| B300 | 288 GB | $3.75spot | Verda | |
| GB300 | 288 GB | $4.31spot | Verda |
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| DeepSeek R1 | $0.200 | $0.400 | 17 |
| DeepSeek V3 | $0.200 | $0.200 | 16 |
| DeepSeek V3.2 | $0.269 | $0.400 | 9 |
| DeepSeek-V3.1 | $0.270 | $1.00 | 9 |
| DeepSeek V4 Flash | $0.050 | $0.101 | 9 |
| DeepSeek V4 Pro | $0.435 | $0.870 | 6 |
FAQ
What is the cheapest API for DeepSeek R1 Distill Llama 70B?
As of 2026-08-23, the lowest input price for DeepSeek R1 Distill Llama 70B is DeepInfra at $0.200 per 1M input tokens. The lowest output price is Nscale at $0.375 per 1M output tokens. The cheapest provider for a real workload depends on its input/output mix. The next-lowest input price is Nebius at $0.250, a 25% difference.
How much VRAM do you need to self-host DeepSeek R1 Distill Llama 70B?
DeepSeek R1 Distill Llama 70B has 70B parameters, so plan for roughly 97 GB of VRAM at FP8 or 49 GB at INT4/AWQ, KV-cache headroom included. A single H200 NVL (141 GB) fits it.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON