MiniMax M2.5 API pricing
6 providers serve MiniMax M2.5. MiniMax M2.5 is an open-weight model (Modified MIT) with 229B total parameters, 10B active per token (mixture-of-experts), up to 1000K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-08-23 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| AWS Bedrock | $0.300 | $1.20 | - | 1000K | Use |
| Baseten | $0.300 | $1.20 | - | - | Use |
| MiniMax | $0.300 | $1.20 | $0.030 | 1000K | Use |
| OpenRouter | $0.300 | $1.10output floor | $0.150 | 197K | Use |
| W&B Inference | $0.300 | $1.20 | - | 197K | Use |
| Tensormesh | $0.300 | $1.20 | - | 197K |
The cheapest way to run MiniMax M2.5
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $54.00/month, at OpenRouter. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, a busy self-hosted deployment can work out cheaper (~$0.137/1M output on 2× MI300X). See the math below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: MiniMax M2.5 self-hosting economics
MiniMax M2.5 is open-weight (Modified MIT), so the API price above competes with the GPU-hour market. 229B total parameters (MoE, ~10B active per token) needs roughly 315 GB VRAM at FP8 or 158 GB at INT4, KV-cache headroom included.
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| MiniMax M2.1 | $0.270 | $1.20 | 6 |
| MiniMax M2 | $0.255 | $1.02 | 6 |
| MiniMax M2.7 | $0.240 | $0.550 | 4 |
| MiniMax M3 | $0.300 | $1.20 | 3 |
| MiniMax M3 (batch) | $0.300 | $1.20 | 1 |
FAQ
What is the cheapest API for MiniMax M2.5?
As of 2026-08-23, the lowest input price for MiniMax M2.5, $0.300 per 1M tokens, is the same list price at AWS Bedrock, Baseten, MiniMax and 3 more. The lowest output price is OpenRouter at $1.10 per 1M output tokens. The cheapest provider for a real workload depends on its input/output mix.
How much VRAM do you need to self-host MiniMax M2.5?
MiniMax M2.5 has 229B parameters, so plan for roughly 315 GB of VRAM at FP8 or 158 GB at INT4/AWQ, KV-cache headroom included. That exceeds a single GPU. A typical node is 2× MI300X.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON