Together AI pricing & review
High-performance open-weight inference (Llama, DeepSeek, Qwen) on a custom stack, plus fine-tuning and GPU clusters. Consistently among the fastest and cheapest for open models.
Where Together AI wins
- Top-tier open-model throughput
- Fine-tuning pipeline
- Dedicated endpoints and clusters
Together AI model pricing
List prices refreshed 2026-10-08 ยท cheapest 25 shown| Model | Input $/1M | Output $/1M | Market floor in |
|---|---|---|---|
| GPT-OSS 20B | $0.050 | $0.200 | $0.015at Darkbloom |
| Llama 3.2 3B | $0.060 | $0.060 | $0.020at DeepInfra |
| Llama 3.2 1B Instruct | $0.060 | $0.060 | $0.020at Novita AI |
| NVIDIA-Nemotron-Nano-9B-V2 | $0.060 | $0.250 | $0.040at DeepInfra |
| Mistral Small 3 | $0.100 | $0.300 | $0.050at DeepInfra |
| DeepSeek V4 Flash | $0.140 | $0.280 | $0.030at OpenRouter |
| GPT-OSS 120B | $0.150 | $0.600 | $0.030at W&B Inference |
| GLM 5.3 Flash | $0.150 | $0.500 | $0.110at Sail |
| Qwen3 Next 80B A3B Instruct | $0.150 | $1.50 | $0.090at DeepInfra |
| Qwen3 Next 80B A3B Thinking | $0.150 | $1.50 | $0.140at DeepInfra |
| Qwen3.8 Flash | $0.150 | $0.470 | $0.113at Aihubmix |
| Qwen3.5-9B | $0.170 | $0.250 | $0.100at DeepInfra |
| Llama 4 Scout | $0.180 | $0.590 | $0.050at Lambda |
| Qwen3 VL 8B Instruct | $0.180 | $0.680 | $0.080at Novita AI |
| DeepSeek-R1-Distill-Qwen-1.5B | $0.180 | $0.180 | $0.090at Nscale |
| Meta-Llama-3.1-8B-Instruct-Turbo | $0.180 | $0.180 | $0.020at DeepInfra |
| GLM-4.5 Air | $0.200 | $1.10 | $0.125at Pinstripes |
| Llama-3-8b | $0.200 | $0.200 | $0.030at DeepInfra |
| Ministral-3-14b-2512 | $0.200 | $0.200 | $0.200at AWS Bedrock |
| Mistral-7B-Instruct-V0.1 | $0.200 | $0.200 | $0.150at Anyscale |
| MiniMax M3 | $0.300 | $1.20 | $0.230at W&B Inference |
| DeepSeek V4.1 Flash | $0.300 | $1.20 | $0.044at OpenRouter |
| MiniMax M2.7 | $0.300 | $1.20 | $0.210at OpenRouter |
| Qwen3.7 Plus | $0.320 | $1.28 | $0.282at Aihubmix |
| Muse Glimmer 30B | $0.350 | $1.50 | $0.300at DeepInfra |