DeepInfra pricing & review
Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options.
Where DeepInfra wins
- Among the cheapest open-model tokens anywhere
- OpenAI-compatible API
- Dedicated GPU deployments
DeepInfra model pricing
List prices refreshed 2026-08-23 ยท cheapest 25 shown| Model | Input $/1M | Output $/1M | Market floor in |
|---|---|---|---|
| Mistral Nemo | $0.020 | $0.040 | $0.019at OpenRouter |
| Llama 3.2 3B | $0.020 | $0.020 | $0.020at DeepInfra |
| Meta-Llama-3.1-8B-Instruct-Turbo | $0.020 | $0.030 | $0.020at DeepInfra |
| Llama 3.1 8B | $0.030 | $0.050 | $0.020at Nebius |
| Llama-3-8b | $0.030 | $0.060 | $0.030at DeepInfra |
| GPT-OSS 20B | $0.040 | $0.150 | $0.015at Darkbloom |
| Qwen2.5-7B-Instruct | $0.040 | $0.100 | $0.040at DeepInfra |
| Gemma 3 4B | $0.040 | $0.080 | $0.040at DeepInfra |
| NVIDIA-Nemotron-Nano-9B-V2 | $0.040 | $0.160 | $0.040at DeepInfra |
| Llama-3.2-11b-Vision-Instruct | $0.049 | $0.049 | $0.049at Cloudflare |
| GPT-OSS 120B | $0.050 | $0.450 | $0.050at DeepInfra |
| Gemma 3 12B | $0.050 | $0.100 | $0.050at DeepInfra |
| Mistral Small 3 | $0.050 | $0.080 | $0.050at DeepInfra |
| Nemotron 3.5 Lightning | $0.050 | $0.200 | $0.050at DeepInfra |
| Llama-Guard-3-8B | $0.055 | $0.055 | $0.020at Nebius |
| Qwen3 14B | $0.060 | $0.240 | $0.060at DeepInfra |
| Phi-4 | $0.070 | $0.140 | $0.070at DeepInfra |
| Mistral Small 3.2 24B | $0.075 | $0.200 | $0.075at DeepInfra |
| Llama 4 Scout | $0.080 | $0.300 | $0.050at Lambda |
| Qwen3 30B A3B | $0.080 | $0.290 | $0.051at Cloudflare |
| MythoMax 13B | $0.080 | $0.090 | $0.080at DeepInfra |
| Gemma 3 27B | $0.090 | $0.160 | $0.060at Nebius |
| Qwen3 32B | $0.100 | $0.280 | $0.050at Lambda |
| Gemini-2.0-Flash-001 | $0.100 | $0.400 | $0.100at DeepInfra |
| Meta-Llama-3.1-70B-Instruct-Turbo | $0.100 | $0.280 | $0.100at DeepInfra |