DeepInfra vs Groq
Same workloads, both price lists, refreshed daily. On shared line items today: DeepInfra is cheaper on 7 of 7 shared models (input or output price), Groq on 0.
DeepInfra
Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options..
- Among the cheapest open-model tokens anywhere
- OpenAI-compatible API
- Dedicated GPU deployments
Groq
Custom LPU silicon serving open models at hundreds of tokens/second: the speed king for latency-sensitive apps, with simple per-token pricing..
- Fastest tokens/sec in the market
- Simple pricing
- Generous free tier for prototyping
Shared models, priced by both
List prices refreshed 2026-08-23 ยท input $/1M tokens| Model | DeepInfra in/out | Groq in/out | Cheaper input |
|---|---|---|---|
| GPT-OSS 20B | $0.040 / $0.150 | $0.075 / $0.300 | DeepInfra |
| GPT-OSS 120B | $0.050 / $0.450 | $0.150 / $0.600 | DeepInfra |
| Llama 4 Scout | $0.080 / $0.300 | $0.110 / $0.340 | DeepInfra |
| Qwen3 32B | $0.100 / $0.280 | $0.290 / $0.590 | DeepInfra |
| Llama 4 Maverick | $0.150 / $0.600 | $0.200 / $0.600 | DeepInfra |
| Llama Guard 4 12B | $0.180 / $0.180 | $0.200 / $0.200 | DeepInfra |
| Kimi K2 | $0.500 / $2.00 | $1.00 / $3.00 | DeepInfra |
Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.