Fireworks AI vs Replicate

Same workloads, both price lists, refreshed daily. On shared line items today: Fireworks AI is cheaper on 6 of 7 shared models (input or output price), Replicate on 0.

Fireworks AI

Speed-focused open-model inference with FireAttention serving stack, function-calling models and image generation. Strong latency SLAs for production apps.

  • Very low latency serving
  • Function-calling optimized
  • Enterprise SLAs

Visit Fireworks AI

Replicate

Run thousands of community models (LLMs, image, audio, video) behind one API with per-second GPU billing: the fastest path from 'found a model on GitHub' to production endpoint..

  • Massive community model zoo
  • Per-second billing, scale to zero
  • Custom models via Cog

Visit Replicate

Shared models, priced by both

List prices refreshed 2026-08-23 ยท input $/1M tokens
Model Fireworks AI in/out Replicate in/out Cheaper input
GPT-OSS 20B $0.070 / $0.300 $0.090 / $0.360 Fireworks AI
GPT-OSS 120B $0.150 / $0.600 $0.180 / $0.720 Fireworks AI
Mistral-7b-Instruct-V0.2 $0.200 / $0.200 $0.050 / $0.250 Replicate
Qwen3 235B A22B $0.220 / $0.880 $0.264 / $1.06 Fireworks AI
DeepSeek-V3.1 $0.560 / $1.68 $0.672 / $2.02 Fireworks AI
DeepSeek V3 $0.900 / $0.900 $1.45 / $1.45 Fireworks AI
DeepSeek R1 $3.00 / $8.00 $3.75 / $10.00 Fireworks AI

Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.