Llama 3.1 8B API pricing
13 providers serve Llama 3.1 8B. Llama 3.1 8B is an open-weight model (Llama 3.1 Community) with 8B total parameters, up to 131K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-10-08 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| Nebius | $0.020 | $0.060 | - | 128K | Use |
| Novita AI | $0.020 | $0.050 | - | 16K | Use |
| DeepInfra | $0.030 | $0.050 | - | 131K | Use |
| Nscale | $0.030 | $0.030output floor | - | - | Use |
| Llamagate | $0.030 | $0.050 | - | 131K | |
| Vercel AI Gateway | $0.050 | $0.080 | - | 131K | Use |
| OpenRouter | $0.050 | $0.080 | $0.025 | 131K | Use |
| OVHcloud | $0.100 | $0.100 | - | 131K | Use |
| Hyperbolic | $0.120 | $0.300 | - | 33K | Use |
| Cloudflare | $0.152 | $0.287 | - | 32K | Use |
| Perplexity | $0.200 | $0.200 | - | 131K | Use |
| W&B Inference | $0.220 | $0.220 | - | 131K | Use |
| Oracle OCI | $0.720 | $0.720 | - | 128K | Use |
The cheapest way to run Llama 3.1 8B
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $2.90/month, at Novita AI. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, the API floor remains below our self-hosting estimate of ~$0.036/1M. Details below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: Llama 3.1 8B self-hosting economics
Llama 3.1 8B is open-weight (Llama 3.1 Community), so the API price above competes with the GPU-hour market. 8B total parameters needs roughly 11 GB VRAM at FP8 or 6 GB at INT4, KV-cache headroom included.
| GPU that fits (FP8, single card) | VRAM | From $/hr | Cheapest at | |
|---|---|---|---|---|
| V100 | 32 GB | $0.055marketplace | Vast.ai | Rent |
| RTX 3090 | 24 GB | $0.121marketplace | Vast.ai | Rent |
| RTX 4090 | 24 GB | $0.136marketplace | Vast.ai | Rent |
| RTX A5000 | 24 GB | $0.160community | RunPod | Rent |
| RTX A4000 | 16 GB | $0.170community | RunPod | Rent |
| RTX 4000 SFF Ada Generation | 20 GB | $0.180community | RunPod | Rent |
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| Llama 3.3 70B | $0.120 | $0.200 | 17 |
| Llama 3.1 70B Instruct | $0.120 | $0.300 | 10 |
| Llama 4 Scout | $0.050 | $0.100 | 9 |
| Llama 3.2 3B | $0.020 | $0.020 | 9 |
| Llama 4 Maverick | $0.050 | $0.100 | 8 |
| Llama-3-70b | $0.120 | $0.300 | 7 |
FAQ
What is the cheapest API for Llama 3.1 8B?
As of 2026-10-08, the lowest input price for Llama 3.1 8B, $0.020 per 1M tokens, is the same list price at Nebius and Novita AI. The lowest output price is Nscale at $0.030 per 1M output tokens. The cheapest provider for a real workload depends on its input/output mix. The next-lowest input price is DeepInfra at $0.030, a 50% difference.
How much VRAM do you need to self-host Llama 3.1 8B?
Llama 3.1 8B has 8B parameters, so plan for roughly 11 GB of VRAM at FP8 or 6 GB at INT4/AWQ, KV-cache headroom included. A single V100 (32 GB) fits it.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON