DeepInfra vs Google Gemini API
Same workloads, both price lists, refreshed daily. List prices tie on all 7 shared line items today, so the choice comes down to platform differences (routing, limits, tooling), summarized below.
DeepInfra
Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options..
- Among the cheapest open-model tokens anywhere
- OpenAI-compatible API
- Dedicated GPU deployments
Google Gemini API
Gemini models with the most generous free tier of any frontier lab and very cheap Flash-class models. Vertex AI offers the same models with enterprise controls.
- Generous free tier
- Cheapest frontier-adjacent small models
- Native multimodal
Shared models, priced by both
List prices refreshed 2026-08-28 ยท input $/1M tokens| Model | DeepInfra in/out | Google Gemini API in/out | Cheaper input |
|---|---|---|---|
| Gemini-2.0-Flash-001 | $0.100 / $0.400 | $0.100 / $0.400 | tie |
| Gemini 3.1 Flash Lite | $0.250 / $1.50 | $0.250 / $1.50 | tie |
| Gemini 2.5 Flash | $0.300 / $2.50 | $0.300 / $2.50 | tie |
| Gemini 3.7 Flash | $0.750 / $3.75 | $0.750 / $3.75 | tie |
| Gemini 2.5 Pro | $1.25 / $10.00 | $1.25 / $10.00 | tie |
| Gemini 3.5 Flash | $1.50 / $9.00 | $1.50 / $9.00 | tie |
| Gemini 3.1 Pro Preview | $2.00 / $12.00 | $2.00 / $12.00 | tie |
Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.