DeepInfra vs Replicate
Same workloads, both price lists, refreshed daily. On shared line items today: DeepInfra is cheaper on 8 of 11 shared models (input or output price), Replicate on 2.
DeepInfra
Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options..
- Among the cheapest open-model tokens anywhere
- OpenAI-compatible API
- Dedicated GPU deployments
Replicate
Run thousands of community models (LLMs, image, audio, video) behind one API with per-second GPU billing: the fastest path from 'found a model on GitHub' to production endpoint..
- Massive community model zoo
- Per-second billing, scale to zero
- Custom models via Cog
Shared models, priced by both
List prices refreshed 2026-08-23 ยท input $/1M tokens| Model | DeepInfra in/out | Replicate in/out | Cheaper input |
|---|---|---|---|
| Llama-3-8b | $0.030 / $0.060 | $0.050 / $0.250 | DeepInfra |
| GPT-OSS 20B | $0.040 / $0.150 | $0.090 / $0.360 | DeepInfra |
| GPT-OSS 120B | $0.050 / $0.450 | $0.180 / $0.720 | DeepInfra |
| Qwen3 235B A22B | $0.180 / $0.540 | $0.264 / $1.06 | DeepInfra |
| DeepSeek-V3.1 | $0.270 / $1.00 | $0.672 / $2.02 | DeepInfra |
| Gemini 2.5 Flash | $0.300 / $2.50 | $2.50 / $2.50 | DeepInfra |
| DeepSeek V3 | $0.380 / $0.890 | $1.45 / $1.45 | DeepInfra |
| Mixtral-8x7B-Instruct-V0.1 | $0.400 / $0.400 | $0.300 / $1.00 | Replicate |
| DeepSeek R1 | $0.700 / $2.40 | $3.75 / $10.00 | DeepInfra |
| Claude-3.7-Sonnet | $3.30 / $16.50 | $3.00 / $15.00 | Replicate |
| Claude Sonnet 4 | $3.30 / $16.50 | $3.00 / $15.00 | Replicate |
Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.