DeepInfra vs Fireworks AI
Same workloads, both price lists, refreshed daily. On shared line items today: DeepInfra is cheaper on 24 of 28 shared models (input or output price), Fireworks AI on 1.
DeepInfra
Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options..
- Among the cheapest open-model tokens anywhere
- OpenAI-compatible API
- Dedicated GPU deployments
Fireworks AI
Speed-focused open-model inference with FireAttention serving stack, function-calling models and image generation. Strong latency SLAs for production apps.
- Very low latency serving
- Function-calling optimized
- Enterprise SLAs
Shared models, priced by both
List prices refreshed 2026-08-23 ยท input $/1M tokens| Model | DeepInfra in/out | Fireworks AI in/out | Cheaper input |
|---|---|---|---|
| Mistral Nemo | $0.020 / $0.040 | $0.200 / $0.200 | DeepInfra |
| GPT-OSS 20B | $0.040 / $0.150 | $0.070 / $0.300 | DeepInfra |
| Qwen2.5-7B-Instruct | $0.040 / $0.100 | $0.200 / $0.200 | DeepInfra |
| NVIDIA-Nemotron-Nano-9B-V2 | $0.040 / $0.160 | $0.200 / $0.200 | DeepInfra |
| GPT-OSS 120B | $0.050 / $0.450 | $0.150 / $0.600 | DeepInfra |
| Mistral Small 3 | $0.050 / $0.080 | $0.900 / $0.900 | DeepInfra |
| Llama-Guard-3-8B | $0.055 / $0.055 | $0.200 / $0.200 | DeepInfra |
| Qwen3 14B | $0.060 / $0.240 | $0.200 / $0.200 | DeepInfra |
| Qwen3 30B A3B | $0.080 / $0.290 | $0.150 / $0.600 | DeepInfra |
| MythoMax 13B | $0.080 / $0.090 | $0.200 / $0.200 | DeepInfra |
| Gemma 3 27B | $0.090 / $0.160 | $0.900 / $0.900 | DeepInfra |
| Qwen3 32B | $0.100 / $0.280 | $0.900 / $0.900 | DeepInfra |
| Qwen2p5-72b | $0.120 / $0.390 | $0.900 / $0.900 | DeepInfra |
| Qwen3 Next 80B A3B Instruct | $0.140 / $1.40 | $0.900 / $0.900 | DeepInfra |
| Qwen3 Next 80B A3B Thinking | $0.140 / $1.40 | $0.900 / $0.900 | DeepInfra |
| QwQ-32B | $0.150 / $0.400 | $0.900 / $0.900 | DeepInfra |
| Qwen3 235B A22B | $0.180 / $0.540 | $0.220 / $0.880 | DeepInfra |
| DeepSeek R1 Distill Llama 70B | $0.200 / $0.600 | $0.900 / $0.900 | DeepInfra |
| Qwen2.5-VL-32B-Instruct | $0.200 / $0.600 | $0.900 / $0.900 | DeepInfra |
| DeepSeek-V3.1 | $0.270 / $1.00 | $0.560 / $1.68 | DeepInfra |
| DeepSeek R1 Distill Qwen 32B | $0.270 / $0.270 | $0.900 / $0.900 | DeepInfra |
| DeepSeek V3.1 Terminus | $0.270 / $1.00 | $0.560 / $1.68 | DeepInfra |
| Qwen3 235B A22B Thinking 2507 | $0.300 / $2.90 | $0.220 / $0.880 | Fireworks AI |
| DeepSeek V3 | $0.380 / $0.890 | $0.900 / $0.900 | DeepInfra |
| Qwen3-Coder-480b-A35b-Instruct | $0.400 / $1.60 | $0.450 / $1.80 | DeepInfra |
| GLM-4.5 | $0.400 / $1.60 | $0.550 / $2.19 | DeepInfra |
| Kimi K2 | $0.500 / $2.00 | $0.600 / $2.50 | DeepInfra |
| DeepSeek R1 | $0.700 / $2.40 | $3.00 / $8.00 | DeepInfra |
Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.