GLM-4.7 API pricing
10 providers serve GLM-4.7. GLM-4.7 is an open-weight model (MIT) with 358B total parameters, 32B active per token (mixture-of-experts), up to 205K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-08-23 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| GMI Cloud | $0.400 | $2.00 | - | 203K | Use |
| OpenRouter | $0.400 | $1.50output floor | - | 203K | Use |
| Together AI | $0.450 | $2.00 | - | 200K | Use |
| Fireworks AI | $0.600 | $2.20 | $0.300 | 203K | Use |
| Baseten | $0.600 | $2.20 | - | - | Use |
| Vertex AI Zai Models | $0.600 | $2.20 | - | 200K | |
| AWS Bedrock | $0.600 | $2.20 | - | 200K | Use |
| Z.ai | $0.600 | $2.20 | $0.110 | 200K | Use |
| Novita AI | $0.600 | $2.20 | $0.110 | 205K | Use |
| Cerebras | $2.25 | $2.75 | - | 128K | Use |
The cheapest way to run GLM-4.7
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $73.00/month, at OpenRouter. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, a busy self-hosted deployment can work out cheaper (~$0.659/1M output on 3× MI300X). See the math below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: GLM-4.7 self-hosting economics
GLM-4.7 is open-weight (MIT), so the API price above competes with the GPU-hour market. 358B total parameters (MoE, ~32B active per token) needs roughly 493 GB VRAM at FP8 or 247 GB at INT4, KV-cache headroom included.
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| GLM-4.6 | $0.400 | $1.75 | 8 |
| GLM-4.5 Air | $0.125 | $0.450 | 7 |
| GLM-5.2 | $0.610 | $1.98 | 6 |
| GLM-4.5 | $0.400 | $1.60 | 6 |
| GLM-5 | $0.800 | $2.56 | 5 |
| GLM-5.1 | $1.05 | $3.50 | 4 |
FAQ
What is the cheapest API for GLM-4.7?
As of 2026-08-23, the lowest input price for GLM-4.7, $0.400 per 1M tokens, is the same list price at GMI Cloud and OpenRouter. The lowest output price is OpenRouter at $1.50 per 1M output tokens. The cheapest provider for a real workload depends on its input/output mix. The next-lowest input price is Together AI at $0.450, a 12% difference.
How much VRAM do you need to self-host GLM-4.7?
GLM-4.7 has 358B parameters, so plan for roughly 493 GB of VRAM at FP8 or 247 GB at INT4/AWQ, KV-cache headroom included. That exceeds a single GPU. A typical node is 3× MI300X.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON