Cerebras pricing & review
Wafer-scale silicon serving open models at extreme speed (2000+ tokens/sec on some models). Competes with Groq for the latency crown.
Where Cerebras wins
- Extreme generation speed
- Simple per-token pricing
Cerebras model pricing
List prices refreshed 2026-08-23 ยท cheapest 25 shown| Model | Input $/1M | Output $/1M | Market floor in |
|---|---|---|---|
| Llama3.1-8b | $0.100 | $0.100 | $0.025at Lambda |
| GPT-OSS 120B | $0.350 | $0.750 | $0.050at DeepInfra |
| Qwen-3-32b | $0.400 | $0.800 | $0.100at Vercel AI Gateway |
| Llama3.1-70b | $0.600 | $0.600 | $0.120at Lambda |
| Llama 3.3 70B | $0.850 | $1.20 | $0.100at OpenRouter |
| GLM-4.7 | $2.25 | $2.75 | $0.400at GMI Cloud |
| GLM-4.6 | $2.25 | $2.75 | $0.400at OpenRouter |