What self-hosting an LLM actually costs

Every open-weight model we track, priced against today's cheapest fitting GPU rental: the VRAM it needs, the hardware plan, an estimated self-hosted $/1M output tokens at 50% utilization, and the cheapest API price it has to beat. At these defaults the API wins for 17 of 51 models. Utilization is the whole game, and the calculator lets you replace every assumption with your own numbers.

GPU floors 2026-08-15 · token floors 2026-08-15
Model Params VRAM FP8 / INT4 Cheapest hardware (FP8) Self-host est.* Cheapest API Verdict at 50% util
Llama 3.2 3B 3.2B 5 / 3 GB V100 $0.027/hr $0.0071/1M $0.020/1M out self-host, 2.8×
Llama 3.1 8B 8B 11 / 6 GB V100 $0.027/hr $0.018/1M $0.030/1M out self-host, 1.7×
IBM Granite 4.1 8B 8.8B 13 / 7 GB V100 $0.027/hr $0.020/1M $0.100/1M out self-host, 5.1×
Gemma 3 12B 12B 17 / 9 GB V100 $0.027/hr $0.027/1M $0.100/1M out self-host, 3.8×
Phi-4 14.7B 21 / 11 GB V100 $0.027/hr $0.033/1M $0.140/1M out self-host, 4.3×
GPT-OSS 20B 21B 3.6B act 29 / 15 GB V100 $0.027/hr $0.0080/1M $0.070/1M out self-host, 8.8×
Mistral Small 3.1 24B 33 / 17 GB A100 PCIe $0.267/hr $0.144/1M - -
Gemma 3 27B 27B 38 / 19 GB A100 PCIe $0.267/hr $0.162/1M $0.160/1M out API, 1.0×
Qwen3.6 27B 27.8B 39 / 20 GB A100 PCIe $0.267/hr $0.167/1M $0.500/1M out self-host, 3.0×
Qwen3.5 27B 27.8B 39 / 20 GB A100 PCIe $0.267/hr $0.167/1M $2.40/1M out self-host, 14.4×
Nemotron 3 Nano 30B 3.5B act 42 / 21 GB A100 PCIe $0.267/hr $0.021/1M $0.200/1M out self-host, 9.5×
IBM Granite 4.1 30B 30B 42 / 21 GB A100 PCIe $0.267/hr $0.180/1M - -
Qwen3 30B A3B 30.5B 3.3B act 42 / 21 GB A100 PCIe $0.267/hr $0.020/1M $0.200/1M out self-host, 10.1×
GLM-4.7 Flash 31B 3B act 43 / 22 GB A100 PCIe $0.267/hr $0.018/1M $0.400/1M out self-host, 22.2×
Qwen3 32B 32.8B 46 / 23 GB A100 PCIe $0.267/hr $0.196/1M $0.100/1M out API, 2.0×
DeepSeek R1 Distill Qwen 32B 32.8B 46 / 23 GB A100 PCIe $0.267/hr $0.196/1M $0.150/1M out API, 1.3×
Qwen3.6 35B A3B 36B 3B act 50 / 25 GB A100 PCIe $0.267/hr $0.018/1M $0.450/1M out self-host, 25.0×
Qwen3.5 35B A3B 36B 3B act 50 / 25 GB A100 PCIe $0.267/hr $0.018/1M $2.00/1M out self-host, 111.2×
Mixtral 8x7B 46.7B 12.9B act 65 / 33 GB A100 PCIe $0.267/hr $0.077/1M $0.280/1M out self-host, 3.6×
Llama 3.3 70B 70B 97 / 49 GB H200 NVL $0.500/hr $0.376/1M $0.200/1M out API, 1.9×
DeepSeek R1 Distill Llama 70B 70B 97 / 49 GB H200 NVL $0.500/hr $0.376/1M $0.375/1M out API, 1.0×
GLM-4.5 Air 106B 12B act 146 / 73 GB MI300X $0.500/hr $0.082/1M $0.450/1M out self-host, 5.5×
Llama 4 Scout 109B 17B act 150 / 75 GB MI300X $0.500/hr $0.117/1M $0.100/1M out API, 1.2×
GPT-OSS 120B 117B 5.1B act 161 / 81 GB MI300X $0.500/hr $0.035/1M $0.250/1M out self-host, 7.1×
Leanstral 1.5 119B 6.5B act 164 / 82 GB MI300X $0.500/hr $0.045/1M - -
Mistral Small 4 119.4B 165 / 83 GB MI300X $0.500/hr $0.819/1M - -
Nemotron 3 Super 120B 12B act 165 / 83 GB MI300X $0.500/hr $0.082/1M $0.400/1M out self-host, 4.9×
Mistral Medium 3.5 128B 176 / 88 GB MI300X $0.500/hr $0.876/1M $7.50/1M out self-host, 8.6×
MiniMax M2.7 229B 10B act 315 / 158 GB MI300X $1.00/hr $0.137/1M $0.550/1M out self-host, 4.0×
MiniMax M2.5 229B 10B act 315 / 158 GB MI300X $1.00/hr $0.137/1M $1.10/1M out self-host, 8.0×
MiniMax M2.1 229B 10B act 315 / 158 GB MI300X $1.00/hr $0.137/1M $1.20/1M out self-host, 8.7×
MiniMax M2 229B 10B act 315 / 158 GB MI300X $1.00/hr $0.137/1M $1.02/1M out self-host, 7.4×
Qwen3 235B A22B 235B 22B act 324 / 162 GB MI300X $1.00/hr $0.302/1M $0.100/1M out API, 3.0×
DeepSeek V4 Flash 284B 13B act 391 / 196 GB H200 NVL $1.50/hr $0.209/1M $0.131/1M out API, 1.6×
GLM-4.5 355B 32B act 489 / 245 GB MI300X $1.50/hr $0.659/1M $1.60/1M out self-host, 2.4×
GLM-4.6 357B 32B act 491 / 246 GB MI300X $1.50/hr $0.659/1M $1.75/1M out self-host, 2.7×
GLM-4.7 358B 32B act 493 / 247 GB MI300X $1.50/hr $0.659/1M $1.50/1M out self-host, 2.3×
Qwen3.5 397B A17B 397B 17B act 546 / 273 GB MI300X $1.50/hr $0.350/1M $3.60/1M out self-host, 10.3×
Llama 4 Maverick 400B 17B act 551 / 276 GB MI300X $1.50/hr $0.350/1M $0.100/1M out API, 3.5×
Llama 3.1 405B 405B 557 / 279 GB MI300X $1.50/hr $6.17/1M $0.300/1M out API, 20.6×
MiniMax M3 428B 23B act 589 / 295 GB MI300X $2.00/hr $0.631/1M $1.20/1M out self-host, 1.9×
Qwen3 Coder 480B 35B act 660 / 330 GB MI300X $2.00/hr $0.960/1M $0.950/1M out API, 1.0×
Nemotron 3 Ultra 550B 55B act 757 / 379 GB MI300X $2.00/hr $1.51/1M $3.60/1M out self-host, 2.4×
DeepSeek R1 685B 37B act 942 / 471 GB MI300X $2.50/hr $1.27/1M $0.400/1M out API, 3.2×
DeepSeek V3 685B 37B act 942 / 471 GB MI300X $2.50/hr $1.27/1M $0.200/1M out API, 6.3×
DeepSeek V3.2 685B 37B act 942 / 471 GB MI300X $2.50/hr $1.27/1M $0.400/1M out API, 3.2×
GLM-5.2 744B 40B act 1024 / 512 GB MI300X $3.00/hr $1.65/1M $1.45/1M out API, 1.1×
GLM-5.1 744B 40B act 1024 / 512 GB MI300X $3.00/hr $1.65/1M $3.50/1M out self-host, 2.1×
GLM-5 744B 40B act 1024 / 512 GB MI300X $3.00/hr $1.65/1M $2.56/1M out self-host, 1.6×
Kimi K2 Thinking 1058B 32B act 1455 / 728 GB MI300X $4.00/hr $1.76/1M $1.20/1M out API, 1.5×
Kimi K2 1058B 32B act 1455 / 728 GB MI300X $4.00/hr $1.76/1M $2.00/1M out self-host, 1.1×
Kimi K2.7 Code 1059B 32B act 1457 / 729 GB MI300X $4.00/hr $1.76/1M $3.50/1M out self-host, 2.0×
Kimi K2.6 1059B 32B act 1457 / 729 GB MI300X $4.00/hr $1.76/1M $2.28/1M out self-host, 1.3×
Kimi K2.5 1059B 32B act 1457 / 729 GB MI300X $4.00/hr $1.76/1M $2.80/1M out self-host, 1.6×
DeepSeek V4 Pro 1600B 49B act 2201 / 1101 GB 12× MI300X $6.00/hr $4.04/1M $0.870/1M out API, 4.6×

*Method: VRAM = params × 1.1 GB/B (FP8) or 0.55 (INT4) × 1.25 headroom. Hardware = cheapest tracked rental fitting that VRAM, multi-GPU nodes rounded up on datacenter cards. Throughput is a conservative planning estimate from active parameters (well-batched vLLM on H100-class), not a benchmark. Self-host $/1M = total $/hr ÷ (tok/s × 3600 × 50% utilization) × 10⁶. API column is today's cheapest listed output price across every provider we track. Input tokens, egress, storage and engineering time are excluded on both sides.

When self-hosting wins

  • Sustained batch workloads that keep utilization above the breakeven line. Run yours through the calculator
  • Fine-tuned or private weights no API serves
  • Data that cannot leave your infrastructure
  • Small models on cheap cards: Llama 3.2 3B runs on a single V100 at $0.027/hr today

Go deeper