Together AI pricing & review

High-performance open-weight inference (Llama, DeepSeek, Qwen) on a custom stack, plus fine-tuning and GPU clusters. Consistently among the fastest and cheapest for open models.

Where Together AI wins

Together AI model pricing

List prices refreshed 2026-10-08 ยท cheapest 25 shown
ModelInput $/1MOutput $/1MMarket floor in
GPT-OSS 20B $0.050 $0.200 $0.015at Darkbloom
Llama 3.2 3B $0.060 $0.060 $0.020at DeepInfra
Llama 3.2 1B Instruct $0.060 $0.060 $0.020at Novita AI
NVIDIA-Nemotron-Nano-9B-V2 $0.060 $0.250 $0.040at DeepInfra
Mistral Small 3 $0.100 $0.300 $0.050at DeepInfra
DeepSeek V4 Flash $0.140 $0.280 $0.030at OpenRouter
GPT-OSS 120B $0.150 $0.600 $0.030at W&B Inference
GLM 5.3 Flash $0.150 $0.500 $0.110at Sail
Qwen3 Next 80B A3B Instruct $0.150 $1.50 $0.090at DeepInfra
Qwen3 Next 80B A3B Thinking $0.150 $1.50 $0.140at DeepInfra
Qwen3.8 Flash $0.150 $0.470 $0.113at Aihubmix
Qwen3.5-9B $0.170 $0.250 $0.100at DeepInfra
Llama 4 Scout $0.180 $0.590 $0.050at Lambda
Qwen3 VL 8B Instruct $0.180 $0.680 $0.080at Novita AI
DeepSeek-R1-Distill-Qwen-1.5B $0.180 $0.180 $0.090at Nscale
Meta-Llama-3.1-8B-Instruct-Turbo $0.180 $0.180 $0.020at DeepInfra
GLM-4.5 Air $0.200 $1.10 $0.125at Pinstripes
Llama-3-8b $0.200 $0.200 $0.030at DeepInfra
Ministral-3-14b-2512 $0.200 $0.200 $0.200at AWS Bedrock
Mistral-7B-Instruct-V0.1 $0.200 $0.200 $0.150at Anyscale
MiniMax M3 $0.300 $1.20 $0.230at W&B Inference
DeepSeek V4.1 Flash $0.300 $1.20 $0.044at OpenRouter
MiniMax M2.7 $0.300 $1.20 $0.210at OpenRouter
Qwen3.7 Plus $0.320 $1.28 $0.282at Aihubmix
Muse Glimmer 30B $0.350 $1.50 $0.300at DeepInfra

Together AI compared