Novita AI vs Replicate

Same workloads, both price lists, refreshed daily. On shared line items today: Novita AI is cheaper on 8 of 8 shared models (input or output price), Replicate on 0.

Novita AI

Combines a serverless LLM API (DeepSeek, Llama, Qwen at aggressive per-token prices) with GPU instances and template deployments, one of the few providers covering both sides of the hosting equation..

  • Both LLM API and GPU rental under one account
  • Aggressive open-weight model pricing
  • Template marketplace for common stacks

Visit Novita AI

Replicate

Run thousands of community models (LLMs, image, audio, video) behind one API with per-second GPU billing: the fastest path from 'found a model on GitHub' to production endpoint..

  • Massive community model zoo
  • Per-second billing, scale to zero
  • Custom models via Cog

Visit Replicate

Shared models, priced by both

List prices refreshed 2026-08-23 ยท input $/1M tokens
Model Novita AI in/out Replicate in/out Cheaper input
GPT-OSS 20B $0.040 / $0.150 $0.090 / $0.360 Novita AI
Llama-3-8b $0.040 / $0.040 $0.050 / $0.250 Novita AI
GPT-OSS 120B $0.050 / $0.250 $0.180 / $0.720 Novita AI
Qwen3 235B A22B $0.200 / $0.800 $0.264 / $1.06 Novita AI
DeepSeek V3 $0.270 / $1.12 $1.45 / $1.45 Novita AI
DeepSeek-V3.1 $0.270 / $1.00 $0.672 / $2.02 Novita AI
Llama-3-70b $0.510 / $0.740 $0.650 / $2.75 Novita AI
DeepSeek R1 $0.700 / $2.50 $3.75 / $10.00 Novita AI

Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.