Fireworks AI vs Novita AI
Same workloads, both price lists, refreshed daily. On shared line items today: Fireworks AI is cheaper on 8 of 30 shared models (input or output price), Novita AI on 16.
Fireworks AI
Speed-focused open-model inference with FireAttention serving stack, function-calling models and image generation. Strong latency SLAs for production apps.
- Very low latency serving
- Function-calling optimized
- Enterprise SLAs
Novita AI
Combines a serverless LLM API (DeepSeek, Llama, Qwen at aggressive per-token prices) with GPU instances and template deployments, one of the few providers covering both sides of the hosting equation..
- Both LLM API and GPU rental under one account
- Aggressive open-weight model pricing
- Template marketplace for common stacks
Shared models, priced by both
List prices refreshed 2026-08-23 ยท input $/1M tokens| Model | Fireworks AI in/out | Novita AI in/out | Cheaper input |
|---|---|---|---|
| GPT-OSS 20B | $0.070 / $0.300 | $0.040 / $0.150 | Novita AI |
| Minimax-M1-80k | $0.100 / $0.100 | $0.550 / $2.20 | Fireworks AI |
| GPT-OSS 120B | $0.150 / $0.600 | $0.050 / $0.250 | Novita AI |
| Qwen3 30B A3B | $0.150 / $0.600 | $0.090 / $0.450 | Novita AI |
| Qwen3 Coder 30B A3B Instruct | $0.150 / $0.600 | $0.070 / $0.270 | Novita AI |
| Qwen3 VL 30B A3B Instruct | $0.150 / $0.600 | $0.200 / $0.700 | Fireworks AI |
| Qwen3 VL 30B A3B Thinking | $0.150 / $0.600 | $0.200 / $1.00 | Fireworks AI |
| Mistral Nemo | $0.200 / $0.200 | $0.040 / $0.170 | Novita AI |
| MythoMax 13B | $0.200 / $0.200 | $0.090 / $0.090 | Novita AI |
| Qwen3 8B | $0.200 / $0.200 | $0.035 / $0.138 | Novita AI |
| Qwen3 VL 8B Instruct | $0.200 / $0.200 | $0.080 / $0.500 | Novita AI |
| Qwen2.5-7B-Instruct | $0.200 / $0.200 | $0.070 / $0.070 | Novita AI |
| DeepSeek-R1-Distill-Qwen-14B | $0.200 / $0.200 | $0.150 / $0.150 | Novita AI |
| Qwen3-4B | $0.200 / $0.200 | $0.030 / $0.030 | Novita AI |
| Qwen3 235B A22B | $0.220 / $0.880 | $0.200 / $0.800 | Novita AI |
| GLM-4.5 Air | $0.220 / $0.880 | $0.130 / $0.850 | Novita AI |
| Qwen3 VL 235B A22B Instruct | $0.220 / $0.880 | $0.300 / $1.50 | Fireworks AI |
| Qwen3 235B A22B Thinking 2507 | $0.220 / $0.880 | $0.300 / $3.00 | Fireworks AI |
| Qwen3 VL 235B A22B Thinking | $0.220 / $0.880 | $0.980 / $3.95 | Fireworks AI |
| MiniMax M2.1 | $0.300 / $1.20 | $0.300 / $1.20 | tie |
| MiniMax M2 | $0.300 / $1.20 | $0.300 / $1.20 | tie |
| Qwen3-Coder-480b-A35b-Instruct | $0.450 / $1.80 | $0.300 / $1.30 | Novita AI |
| GLM-4.6 | $0.550 / $2.19 | $0.550 / $2.20 | tie |
| GLM-4.5 | $0.550 / $2.19 | $0.600 / $2.20 | Fireworks AI |
| DeepSeek V3.2 | $0.560 / $1.68 | $0.269 / $0.400 | Novita AI |
| DeepSeek-V3.1 | $0.560 / $1.68 | $0.270 / $1.00 | Novita AI |
| DeepSeek V3.1 Terminus | $0.560 / $1.68 | $0.270 / $1.00 | Novita AI |
| Kimi K2 | $0.600 / $2.50 | $0.600 / $2.50 | tie |
| GLM-4.7 | $0.600 / $2.20 | $0.600 / $2.20 | tie |
| Kimi K2 Thinking | $0.600 / $2.50 | $0.600 / $2.50 | tie |
Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.