What self-hosting an LLM actually costs
Every open-weight model we track, priced against today's cheapest fitting GPU rental: the VRAM it needs, the hardware plan, an estimated self-hosted $/1M output tokens at 50% utilization, and the cheapest API price it has to beat. At these defaults the API wins for 17 of 51 models. Utilization is the whole game, and the calculator lets you replace every assumption with your own numbers.
GPU floors 2026-08-15 · token floors 2026-08-15| Model | Params | VRAM FP8 / INT4 | Cheapest hardware (FP8) | Self-host est.* | Cheapest API | Verdict at 50% util |
|---|---|---|---|---|---|---|
| Llama 3.2 3B | 3.2B | 5 / 3 GB | 1× V100 $0.027/hr | $0.0071/1M | $0.020/1M out | self-host, 2.8× |
| Llama 3.1 8B | 8B | 11 / 6 GB | 1× V100 $0.027/hr | $0.018/1M | $0.030/1M out | self-host, 1.7× |
| IBM Granite 4.1 8B | 8.8B | 13 / 7 GB | 1× V100 $0.027/hr | $0.020/1M | $0.100/1M out | self-host, 5.1× |
| Gemma 3 12B | 12B | 17 / 9 GB | 1× V100 $0.027/hr | $0.027/1M | $0.100/1M out | self-host, 3.8× |
| Phi-4 | 14.7B | 21 / 11 GB | 1× V100 $0.027/hr | $0.033/1M | $0.140/1M out | self-host, 4.3× |
| GPT-OSS 20B | 21B 3.6B act | 29 / 15 GB | 1× V100 $0.027/hr | $0.0080/1M | $0.070/1M out | self-host, 8.8× |
| Mistral Small 3.1 | 24B | 33 / 17 GB | 1× A100 PCIe $0.267/hr | $0.144/1M | - | - |
| Gemma 3 27B | 27B | 38 / 19 GB | 1× A100 PCIe $0.267/hr | $0.162/1M | $0.160/1M out | API, 1.0× |
| Qwen3.6 27B | 27.8B | 39 / 20 GB | 1× A100 PCIe $0.267/hr | $0.167/1M | $0.500/1M out | self-host, 3.0× |
| Qwen3.5 27B | 27.8B | 39 / 20 GB | 1× A100 PCIe $0.267/hr | $0.167/1M | $2.40/1M out | self-host, 14.4× |
| Nemotron 3 Nano | 30B 3.5B act | 42 / 21 GB | 1× A100 PCIe $0.267/hr | $0.021/1M | $0.200/1M out | self-host, 9.5× |
| IBM Granite 4.1 30B | 30B | 42 / 21 GB | 1× A100 PCIe $0.267/hr | $0.180/1M | - | - |
| Qwen3 30B A3B | 30.5B 3.3B act | 42 / 21 GB | 1× A100 PCIe $0.267/hr | $0.020/1M | $0.200/1M out | self-host, 10.1× |
| GLM-4.7 Flash | 31B 3B act | 43 / 22 GB | 1× A100 PCIe $0.267/hr | $0.018/1M | $0.400/1M out | self-host, 22.2× |
| Qwen3 32B | 32.8B | 46 / 23 GB | 1× A100 PCIe $0.267/hr | $0.196/1M | $0.100/1M out | API, 2.0× |
| DeepSeek R1 Distill Qwen 32B | 32.8B | 46 / 23 GB | 1× A100 PCIe $0.267/hr | $0.196/1M | $0.150/1M out | API, 1.3× |
| Qwen3.6 35B A3B | 36B 3B act | 50 / 25 GB | 1× A100 PCIe $0.267/hr | $0.018/1M | $0.450/1M out | self-host, 25.0× |
| Qwen3.5 35B A3B | 36B 3B act | 50 / 25 GB | 1× A100 PCIe $0.267/hr | $0.018/1M | $2.00/1M out | self-host, 111.2× |
| Mixtral 8x7B | 46.7B 12.9B act | 65 / 33 GB | 1× A100 PCIe $0.267/hr | $0.077/1M | $0.280/1M out | self-host, 3.6× |
| Llama 3.3 70B | 70B | 97 / 49 GB | 1× H200 NVL $0.500/hr | $0.376/1M | $0.200/1M out | API, 1.9× |
| DeepSeek R1 Distill Llama 70B | 70B | 97 / 49 GB | 1× H200 NVL $0.500/hr | $0.376/1M | $0.375/1M out | API, 1.0× |
| GLM-4.5 Air | 106B 12B act | 146 / 73 GB | 1× MI300X $0.500/hr | $0.082/1M | $0.450/1M out | self-host, 5.5× |
| Llama 4 Scout | 109B 17B act | 150 / 75 GB | 1× MI300X $0.500/hr | $0.117/1M | $0.100/1M out | API, 1.2× |
| GPT-OSS 120B | 117B 5.1B act | 161 / 81 GB | 1× MI300X $0.500/hr | $0.035/1M | $0.250/1M out | self-host, 7.1× |
| Leanstral 1.5 | 119B 6.5B act | 164 / 82 GB | 1× MI300X $0.500/hr | $0.045/1M | - | - |
| Mistral Small 4 | 119.4B | 165 / 83 GB | 1× MI300X $0.500/hr | $0.819/1M | - | - |
| Nemotron 3 Super | 120B 12B act | 165 / 83 GB | 1× MI300X $0.500/hr | $0.082/1M | $0.400/1M out | self-host, 4.9× |
| Mistral Medium 3.5 | 128B | 176 / 88 GB | 1× MI300X $0.500/hr | $0.876/1M | $7.50/1M out | self-host, 8.6× |
| MiniMax M2.7 | 229B 10B act | 315 / 158 GB | 2× MI300X $1.00/hr | $0.137/1M | $0.550/1M out | self-host, 4.0× |
| MiniMax M2.5 | 229B 10B act | 315 / 158 GB | 2× MI300X $1.00/hr | $0.137/1M | $1.10/1M out | self-host, 8.0× |
| MiniMax M2.1 | 229B 10B act | 315 / 158 GB | 2× MI300X $1.00/hr | $0.137/1M | $1.20/1M out | self-host, 8.7× |
| MiniMax M2 | 229B 10B act | 315 / 158 GB | 2× MI300X $1.00/hr | $0.137/1M | $1.02/1M out | self-host, 7.4× |
| Qwen3 235B A22B | 235B 22B act | 324 / 162 GB | 2× MI300X $1.00/hr | $0.302/1M | $0.100/1M out | API, 3.0× |
| DeepSeek V4 Flash | 284B 13B act | 391 / 196 GB | 3× H200 NVL $1.50/hr | $0.209/1M | $0.131/1M out | API, 1.6× |
| GLM-4.5 | 355B 32B act | 489 / 245 GB | 3× MI300X $1.50/hr | $0.659/1M | $1.60/1M out | self-host, 2.4× |
| GLM-4.6 | 357B 32B act | 491 / 246 GB | 3× MI300X $1.50/hr | $0.659/1M | $1.75/1M out | self-host, 2.7× |
| GLM-4.7 | 358B 32B act | 493 / 247 GB | 3× MI300X $1.50/hr | $0.659/1M | $1.50/1M out | self-host, 2.3× |
| Qwen3.5 397B A17B | 397B 17B act | 546 / 273 GB | 3× MI300X $1.50/hr | $0.350/1M | $3.60/1M out | self-host, 10.3× |
| Llama 4 Maverick | 400B 17B act | 551 / 276 GB | 3× MI300X $1.50/hr | $0.350/1M | $0.100/1M out | API, 3.5× |
| Llama 3.1 405B | 405B | 557 / 279 GB | 3× MI300X $1.50/hr | $6.17/1M | $0.300/1M out | API, 20.6× |
| MiniMax M3 | 428B 23B act | 589 / 295 GB | 4× MI300X $2.00/hr | $0.631/1M | $1.20/1M out | self-host, 1.9× |
| Qwen3 Coder | 480B 35B act | 660 / 330 GB | 4× MI300X $2.00/hr | $0.960/1M | $0.950/1M out | API, 1.0× |
| Nemotron 3 Ultra | 550B 55B act | 757 / 379 GB | 4× MI300X $2.00/hr | $1.51/1M | $3.60/1M out | self-host, 2.4× |
| DeepSeek R1 | 685B 37B act | 942 / 471 GB | 5× MI300X $2.50/hr | $1.27/1M | $0.400/1M out | API, 3.2× |
| DeepSeek V3 | 685B 37B act | 942 / 471 GB | 5× MI300X $2.50/hr | $1.27/1M | $0.200/1M out | API, 6.3× |
| DeepSeek V3.2 | 685B 37B act | 942 / 471 GB | 5× MI300X $2.50/hr | $1.27/1M | $0.400/1M out | API, 3.2× |
| GLM-5.2 | 744B 40B act | 1024 / 512 GB | 6× MI300X $3.00/hr | $1.65/1M | $1.45/1M out | API, 1.1× |
| GLM-5.1 | 744B 40B act | 1024 / 512 GB | 6× MI300X $3.00/hr | $1.65/1M | $3.50/1M out | self-host, 2.1× |
| GLM-5 | 744B 40B act | 1024 / 512 GB | 6× MI300X $3.00/hr | $1.65/1M | $2.56/1M out | self-host, 1.6× |
| Kimi K2 Thinking | 1058B 32B act | 1455 / 728 GB | 8× MI300X $4.00/hr | $1.76/1M | $1.20/1M out | API, 1.5× |
| Kimi K2 | 1058B 32B act | 1455 / 728 GB | 8× MI300X $4.00/hr | $1.76/1M | $2.00/1M out | self-host, 1.1× |
| Kimi K2.7 Code | 1059B 32B act | 1457 / 729 GB | 8× MI300X $4.00/hr | $1.76/1M | $3.50/1M out | self-host, 2.0× |
| Kimi K2.6 | 1059B 32B act | 1457 / 729 GB | 8× MI300X $4.00/hr | $1.76/1M | $2.28/1M out | self-host, 1.3× |
| Kimi K2.5 | 1059B 32B act | 1457 / 729 GB | 8× MI300X $4.00/hr | $1.76/1M | $2.80/1M out | self-host, 1.6× |
| DeepSeek V4 Pro | 1600B 49B act | 2201 / 1101 GB | 12× MI300X $6.00/hr | $4.04/1M | $0.870/1M out | API, 4.6× |
*Method: VRAM = params × 1.1 GB/B (FP8) or 0.55 (INT4) × 1.25 headroom. Hardware = cheapest tracked rental fitting that VRAM, multi-GPU nodes rounded up on datacenter cards. Throughput is a conservative planning estimate from active parameters (well-batched vLLM on H100-class), not a benchmark. Self-host $/1M = total $/hr ÷ (tok/s × 3600 × 50% utilization) × 10⁶. API column is today's cheapest listed output price across every provider we track. Input tokens, egress, storage and engineering time are excluded on both sides.
When self-hosting wins
- Sustained batch workloads that keep utilization above the breakeven line. Run yours through the calculator
- Fine-tuned or private weights no API serves
- Data that cannot leave your infrastructure
- Small models on cheap cards: Llama 3.2 3B runs on a single V100 at $0.027/hr today
Go deeper
- The API-vs-GPU breakeven math, worked through
- Why the same GPU rents at three different prices
- Every GPU we track, with live floors
- GPU price history: floors move, timing matters
- Serving-stack hosting guides (vLLM, SGLang, Ollama)
- Buying the box instead: owned-hardware breakeven
- Hand your workload to us: ranked plan across API, rented GPUs, and owned hardware