Replicate vs Together AI
Same workloads, both price lists, refreshed daily. On shared line items today: Replicate is cheaper on 0 of 6 shared models (input or output price), Together AI on 5.
Replicate
Run thousands of community models (LLMs, image, audio, video) behind one API with per-second GPU billing: the fastest path from 'found a model on GitHub' to production endpoint..
- Massive community model zoo
- Per-second billing, scale to zero
- Custom models via Cog
Together AI
High-performance open-weight inference (Llama, DeepSeek, Qwen) on a custom stack, plus fine-tuning and GPU clusters. Consistently among the fastest and cheapest for open models.
- Top-tier open-model throughput
- Fine-tuning pipeline
- Dedicated endpoints and clusters
Shared models, priced by both
List prices refreshed 2026-08-23 ยท input $/1M tokens| Model | Replicate in/out | Together AI in/out | Cheaper input |
|---|---|---|---|
| GPT-OSS 20B | $0.090 / $0.360 | $0.050 / $0.200 | Together AI |
| GPT-OSS 120B | $0.180 / $0.720 | $0.150 / $0.600 | Together AI |
| Mixtral-8x7B-Instruct-V0.1 | $0.300 / $1.00 | $0.600 / $0.600 | Replicate |
| DeepSeek-V3.1 | $0.672 / $2.02 | $0.600 / $1.70 | Together AI |
| DeepSeek V3 | $1.45 / $1.45 | $1.25 / $1.25 | Together AI |
| DeepSeek R1 | $3.75 / $10.00 | $3.00 / $7.00 | Together AI |
Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.