DeepInfra vs Google Vertex AI
Same workloads, both price lists, refreshed daily. On shared line items today: DeepInfra is cheaper on 8 of 17 shared models (input or output price), Google Vertex AI on 3.
DeepInfra
Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options..
- Among the cheapest open-model tokens anywhere
- OpenAI-compatible API
- Dedicated GPU deployments
Google Vertex AI
Gemini plus 100+ models in Google Cloud's ML platform, with enterprise governance, grounding and tuning pipelines..
- Enterprise governance
- Grounding with Google Search
- Model Garden variety
Shared models, priced by both
List prices refreshed 2026-08-23 ยท input $/1M tokens| Model | DeepInfra in/out | Google Vertex AI in/out | Cheaper input |
|---|---|---|---|
| Mistral Nemo | $0.020 / $0.040 | $3.00 / $3.00 | DeepInfra |
| GPT-OSS 20B | $0.040 / $0.150 | $0.075 / $0.300 | DeepInfra |
| GPT-OSS 120B | $0.050 / $0.450 | $0.090 / $0.360 | DeepInfra |
| Llama 4 Scout | $0.080 / $0.300 | $0.250 / $0.700 | DeepInfra |
| Gemini-2.0-Flash-001 | $0.100 / $0.400 | $0.150 / $0.600 | DeepInfra |
| Qwen3 Next 80B A3B Instruct | $0.140 / $1.40 | $0.150 / $1.20 | DeepInfra |
| Qwen3 Next 80B A3B Thinking | $0.140 / $1.40 | $0.150 / $1.20 | DeepInfra |
| Llama 4 Maverick | $0.150 / $0.600 | $0.350 / $1.15 | DeepInfra |
| Qwen3 235B A22B | $0.180 / $0.540 | $0.220 / $0.880 | DeepInfra |
| DeepSeek-V3.1 | $0.270 / $1.00 | $0.600 / $1.70 | DeepInfra |
| Gemini 2.5 Flash | $0.300 / $2.50 | $0.300 / $2.50 | tie |
| Qwen3-Coder-480b-A35b-Instruct | $0.400 / $1.60 | $0.220 / $1.80 | Google Vertex AI |
| DeepSeek R1 | $0.700 / $2.40 | $1.35 / $5.40 | DeepInfra |
| Gemini 2.5 Pro | $1.25 / $10.00 | $1.25 / $10.00 | tie |
| Claude-3.7-Sonnet | $3.30 / $16.50 | $3.00 / $15.00 | Google Vertex AI |
| Claude Sonnet 4 | $3.30 / $16.50 | $3.00 / $15.00 | Google Vertex AI |
| Claude Opus 4 | $16.50 / $82.50 | $15.00 / $75.00 | Google Vertex AI |
Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.