DeepInfra vs Replicate

Same workloads, both price lists, refreshed daily. On shared line items today: DeepInfra is cheaper on 8 of 11 shared models (input or output price), Replicate on 2.

DeepInfra

Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options..

  • Among the cheapest open-model tokens anywhere
  • OpenAI-compatible API
  • Dedicated GPU deployments

Visit DeepInfra

Replicate

Run thousands of community models (LLMs, image, audio, video) behind one API with per-second GPU billing: the fastest path from 'found a model on GitHub' to production endpoint..

  • Massive community model zoo
  • Per-second billing, scale to zero
  • Custom models via Cog

Visit Replicate

Shared models, priced by both

List prices refreshed 2026-08-23 ยท input $/1M tokens
Model DeepInfra in/out Replicate in/out Cheaper input
Llama-3-8b $0.030 / $0.060 $0.050 / $0.250 DeepInfra
GPT-OSS 20B $0.040 / $0.150 $0.090 / $0.360 DeepInfra
GPT-OSS 120B $0.050 / $0.450 $0.180 / $0.720 DeepInfra
Qwen3 235B A22B $0.180 / $0.540 $0.264 / $1.06 DeepInfra
DeepSeek-V3.1 $0.270 / $1.00 $0.672 / $2.02 DeepInfra
Gemini 2.5 Flash $0.300 / $2.50 $2.50 / $2.50 DeepInfra
DeepSeek V3 $0.380 / $0.890 $1.45 / $1.45 DeepInfra
Mixtral-8x7B-Instruct-V0.1 $0.400 / $0.400 $0.300 / $1.00 Replicate
DeepSeek R1 $0.700 / $2.40 $3.75 / $10.00 DeepInfra
Claude-3.7-Sonnet $3.30 / $16.50 $3.00 / $15.00 Replicate
Claude Sonnet 4 $3.30 / $16.50 $3.00 / $15.00 Replicate

Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.