Replicate vs Together AI

Same workloads, both price lists, refreshed daily. On shared line items today: Replicate is cheaper on 0 of 6 shared models (input or output price), Together AI on 5.

Replicate

Run thousands of community models (LLMs, image, audio, video) behind one API with per-second GPU billing: the fastest path from 'found a model on GitHub' to production endpoint..

  • Massive community model zoo
  • Per-second billing, scale to zero
  • Custom models via Cog

Visit Replicate

Together AI

High-performance open-weight inference (Llama, DeepSeek, Qwen) on a custom stack, plus fine-tuning and GPU clusters. Consistently among the fastest and cheapest for open models.

  • Top-tier open-model throughput
  • Fine-tuning pipeline
  • Dedicated endpoints and clusters

Visit Together AI

Shared models, priced by both

List prices refreshed 2026-08-23 ยท input $/1M tokens
Model Replicate in/out Together AI in/out Cheaper input
GPT-OSS 20B $0.090 / $0.360 $0.050 / $0.200 Together AI
GPT-OSS 120B $0.180 / $0.720 $0.150 / $0.600 Together AI
Mixtral-8x7B-Instruct-V0.1 $0.300 / $1.00 $0.600 / $0.600 Replicate
DeepSeek-V3.1 $0.672 / $2.02 $0.600 / $1.70 Together AI
DeepSeek V3 $1.45 / $1.45 $1.25 / $1.25 Together AI
DeepSeek R1 $3.75 / $10.00 $3.00 / $7.00 Together AI

Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.