Groq vs Together AI
Same workloads, both price lists, refreshed daily. On shared line items today: Groq is cheaper on 2 of 5 shared models (input or output price), Together AI on 1.
Groq
Custom LPU silicon serving open models at hundreds of tokens/second: the speed king for latency-sensitive apps, with simple per-token pricing..
- Fastest tokens/sec in the market
- Simple pricing
- Generous free tier for prototyping
Together AI
High-performance open-weight inference (Llama, DeepSeek, Qwen) on a custom stack, plus fine-tuning and GPU clusters. Consistently among the fastest and cheapest for open models.
- Top-tier open-model throughput
- Fine-tuning pipeline
- Dedicated endpoints and clusters
Shared models, priced by both
List prices refreshed 2026-08-23 ยท input $/1M tokens| Model | Groq in/out | Together AI in/out | Cheaper input |
|---|---|---|---|
| GPT-OSS 20B | $0.075 / $0.300 | $0.050 / $0.200 | Together AI |
| Llama 4 Scout | $0.110 / $0.340 | $0.180 / $0.590 | Groq |
| GPT-OSS 120B | $0.150 / $0.600 | $0.150 / $0.600 | tie |
| Llama 4 Maverick | $0.200 / $0.600 | $0.270 / $0.850 | Groq |
| Kimi K2 | $1.00 / $3.00 | $1.00 / $3.00 | tie |
Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.