Llama 4 Maverick API pricing
9 providers serve Llama 4 Maverick. Llama 4 Maverick is an open-weight model (Llama 4 Community) with 400B total parameters, 17B active per token (mixture-of-experts), up to 1049K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-08-23 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| Lambda | $0.050 | $0.100output floor | - | 131K | Use |
| DeepInfra | $0.150 | $0.600 | - | 1049K | Use |
| Groq | $0.200 | $0.600 | - | 131K | Use |
| Together AI | $0.270 | $0.850 | - | - | Use |
| Novita AI | $0.270 | $0.850 | - | 1049K | Use |
| Google Vertex AI | $0.350 | $1.15 | - | 1000K | Use |
| SambaNova | $0.630 | $1.80 | - | 131K | Use |
| Oracle OCI | $0.720 | $0.720 | - | 1049K | Use |
| Azure AI Foundry | $1.41 | $0.350 | - | 1000K | Use |
The cheapest way to run Llama 4 Maverick
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $6.50/month, at Lambda. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, the API floor remains below our self-hosting estimate of ~$0.350/1M. Details below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: Llama 4 Maverick self-hosting economics
Llama 4 Maverick is open-weight (Llama 4 Community), so the API price above competes with the GPU-hour market. 400B total parameters (MoE, ~17B active per token) needs roughly 551 GB VRAM at FP8 or 276 GB at INT4, KV-cache headroom included.
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| Llama 3.3 70B | $0.100 | $0.200 | 15 |
| Llama 3.1 8B | $0.020 | $0.030 | 15 |
| Llama 4 Scout | $0.050 | $0.100 | 11 |
| Llama 3.1 70B Instruct | $0.120 | $0.300 | 10 |
| Llama 3.2 3B | $0.020 | $0.020 | 9 |
| Llama-3-70b | $0.120 | $0.300 | 7 |
FAQ
What is the cheapest API for Llama 4 Maverick?
As of 2026-08-23, the lowest input price for Llama 4 Maverick is Lambda at $0.050 per 1M input tokens. The lowest output price is Lambda at $0.100 per 1M output tokens. The cheapest provider for a real workload depends on its input/output mix. The next-lowest input price is DeepInfra at $0.150, a 200% difference.
How much VRAM do you need to self-host Llama 4 Maverick?
Llama 4 Maverick has 400B parameters, so plan for roughly 551 GB of VRAM at FP8 or 276 GB at INT4/AWQ, KV-cache headroom included. That exceeds a single GPU. A typical node is 3× MI300X.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON