Kimi K2 Thinking API pricing
9 providers serve Kimi K2 Thinking. Kimi K2 Thinking is an open-weight model (Modified MIT) with 1058B total parameters, 32B active per token (mixture-of-experts), up to 262K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-08-23 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| AWS Bedrock | $0.600 | $2.50 | - | 128K | Use |
| Fireworks AI | $0.600 | $2.50 | - | 262K | Use |
| Baseten | $0.600 | $2.50 | - | - | Use |
| Moonshot | $0.600 | $2.50 | $0.150 | 262K | Use |
| Vertex AI Moonshot Models | $0.600 | $2.50 | - | 256K | |
| Novita AI | $0.600 | $2.50 | - | 262K | Use |
| OpenRouter marketplace quote | $0.600 | $2.50 | $0.150 | 262K | Use |
| GMI Cloud | $0.800 | $1.20output floor | - | 262K | Use |
| Crusoe | $2.50 | $2.50 | - | 262K | Use |
The cheapest way to run Kimi K2 Thinking
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $92.00/month, at GMI Cloud. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, the API floor remains below our self-hosting estimate of ~$1.76/1M. Details below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: Kimi K2 Thinking self-hosting economics
Kimi K2 Thinking is open-weight (Modified MIT), so the API price above competes with the GPU-hour market. 1058B total parameters (MoE, ~32B active per token) needs roughly 1455 GB VRAM at FP8 or 728 GB at INT4, KV-cache headroom included.
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| Kimi K2 | $0.500 | $2.00 | 11 |
| Kimi K2.5 | $0.500 | $2.80 | 8 |
| Kimi K2.6 | $0.950 | $4.00 | 6 |
| Kimi K2.7 Code | $0.670 | $3.40 | 4 |
| Kimi K3 | $3.00 | $15.00 | 3 |
| MoonshotAI Kimi Latest | $2.00 | $5.00 | 2 |
FAQ
What is the cheapest API for Kimi K2 Thinking?
As of 2026-08-23, the lowest input price for Kimi K2 Thinking, $0.600 per 1M tokens, is the same list price at AWS Bedrock, Fireworks AI, Baseten and 4 more. The lowest output price is GMI Cloud at $1.20 per 1M output tokens. The cheapest provider for a real workload depends on its input/output mix. The next-lowest input price is GMI Cloud at $0.800, a 33% difference.
How much VRAM do you need to self-host Kimi K2 Thinking?
Kimi K2 Thinking has 1058B parameters, so plan for roughly 1455 GB of VRAM at FP8 or 728 GB at INT4/AWQ, KV-cache headroom included. That exceeds a single GPU. A typical node is 8× MI300X.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON