Fireworks AI pricing & review
Speed-focused open-model inference with FireAttention serving stack, function-calling models and image generation. Strong latency SLAs for production apps.
Where Fireworks AI wins
- Very low latency serving
- Function-calling optimized
- Enterprise SLAs
Fireworks AI model pricing
List prices refreshed 2026-08-23 ยท cheapest 25 shown| Model | Input $/1M | Output $/1M | Market floor in |
|---|---|---|---|
| GPT-OSS 20B | $0.070 | $0.300 | $0.015at Darkbloom |
| Ministral-3-3b-2512 | $0.100 | $0.100 | $0.100at AWS Bedrock |
| Phi-3-Mini-128k-Instruct | $0.100 | $0.100 | $0.100at Fireworks AI |
| Qwen2p5-Coder-3b | $0.100 | $0.100 | $0.010at Nscale |
| DeepSeek-R1-Distill-Qwen-1.5B | $0.100 | $0.100 | $0.090at Nscale |
| Minimax-M1-80k | $0.100 | $0.100 | $0.100at Fireworks AI |
| DeepSeek V4 Flash | $0.140 | $0.280 | $0.050at OpenRouter |
| GPT-OSS 120B | $0.150 | $0.600 | $0.050at DeepInfra |
| Qwen3 30B A3B | $0.150 | $0.600 | $0.051at Cloudflare |
| Qwen3 Coder 30B A3B Instruct | $0.150 | $0.600 | $0.070at Novita AI |
| Qwen3 VL 30B A3B Instruct | $0.150 | $0.600 | $0.130at OpenRouter |
| Qwen3 VL 30B A3B Thinking | $0.150 | $0.600 | $0.150at Fireworks AI |
| Mistral Nemo | $0.200 | $0.200 | $0.019at OpenRouter |
| Llama-Guard-3-8B | $0.200 | $0.200 | $0.020at Nebius |
| Mistral-7b | $0.200 | $0.200 | $0.070at Perplexity |
| MythoMax 13B | $0.200 | $0.200 | $0.080at DeepInfra |
| Qwen3 14B | $0.200 | $0.200 | $0.060at DeepInfra |
| Qwen2.5-Coder-7B | $0.200 | $0.200 | $0.010at Nscale |
| Qwen3 8B | $0.200 | $0.200 | $0.035at Novita AI |
| Qwen3 VL 8B Instruct | $0.200 | $0.200 | $0.080at Novita AI |
| Gemma-7b-It | $0.200 | $0.200 | $0.050at Groq |
| Qwen2.5-7B-Instruct | $0.200 | $0.200 | $0.040at DeepInfra |
| Ministral-3-14b-2512 | $0.200 | $0.200 | $0.200at AWS Bedrock |
| Ministral-3-8b-2512 | $0.200 | $0.200 | $0.150at AWS Bedrock |
| DeepSeek-R1-Distill-Qwen-14B | $0.200 | $0.200 | $0.070at Nscale |