Together AI pricing & review
High-performance open-weight inference (Llama, DeepSeek, Qwen) on a custom stack, plus fine-tuning and GPU clusters. Consistently among the fastest and cheapest for open models.
Where Together AI wins
- Top-tier open-model throughput
- Fine-tuning pipeline
- Dedicated endpoints and clusters
Together AI model pricing
List prices refreshed 2026-08-23 ยท cheapest 25 shown| Model | Input $/1M | Output $/1M | Market floor in |
|---|---|---|---|
| GPT-OSS 20B | $0.050 | $0.200 | $0.015at Darkbloom |
| GPT-OSS 120B | $0.150 | $0.600 | $0.050at DeepInfra |
| Qwen3 Next 80B A3B Instruct | $0.150 | $1.50 | $0.100at OpenRouter |
| Qwen3 Next 80B A3B Thinking | $0.150 | $1.50 | $0.140at DeepInfra |
| Llama 4 Scout | $0.180 | $0.590 | $0.050at Lambda |
| Meta-Llama-3.1-8B-Instruct-Turbo | $0.180 | $0.180 | $0.020at DeepInfra |
| GLM-4.5 Air | $0.200 | $1.10 | $0.125at Pinstripes |
| Llama 4 Maverick | $0.270 | $0.850 | $0.050at Lambda |
| GLM-4.7 | $0.450 | $2.00 | $0.400at GMI Cloud |
| Kimi K2.5 | $0.500 | $2.80 | $0.500at Together AI |
| DeepSeek-V3.1 | $0.600 | $1.70 | $0.270at DeepInfra |
| GLM-4.6 | $0.600 | $2.20 | $0.400at OpenRouter |
| Mixtral-8x7B-Instruct-V0.1 | $0.600 | $0.600 | $0.150at Anyscale |
| Qwen3.5 397B A17B | $0.600 | $3.60 | $0.600at OpenRouter |
| Qwen3 235B A22B Thinking 2507 | $0.650 | $3.00 | $0.110at OpenRouter |
| Llama-3.3-70B-Instruct-Turbo | $0.880 | $0.880 | $0.130at DeepInfra |
| Meta-Llama-3.1-70B-Instruct-Turbo | $0.880 | $0.880 | $0.100at DeepInfra |
| Kimi K2 | $1.00 | $3.00 | $0.500at DeepInfra |
| DeepSeek V3 | $1.25 | $1.25 | $0.200at Hyperbolic |
| Qwen3-Coder-480b-A35b-Instruct | $2.00 | $2.00 | $0.220at Google Vertex AI |
| DeepSeek R1 | $3.00 | $7.00 | $0.200at Lambda |