DeepInfra vs Together AI

Same workloads, both price lists, refreshed daily. On shared line items today: DeepInfra is cheaper on 16 of 16 shared models (input or output price), Together AI on 0.

DeepInfra

Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options..

  • Among the cheapest open-model tokens anywhere
  • OpenAI-compatible API
  • Dedicated GPU deployments

Visit DeepInfra

Together AI

High-performance open-weight inference (Llama, DeepSeek, Qwen) on a custom stack, plus fine-tuning and GPU clusters. Consistently among the fastest and cheapest for open models.

  • Top-tier open-model throughput
  • Fine-tuning pipeline
  • Dedicated endpoints and clusters

Visit Together AI

Shared models, priced by both

List prices refreshed 2026-08-23 ยท input $/1M tokens
Model DeepInfra in/out Together AI in/out Cheaper input
Meta-Llama-3.1-8B-Instruct-Turbo $0.020 / $0.030 $0.180 / $0.180 DeepInfra
GPT-OSS 20B $0.040 / $0.150 $0.050 / $0.200 DeepInfra
GPT-OSS 120B $0.050 / $0.450 $0.150 / $0.600 DeepInfra
Llama 4 Scout $0.080 / $0.300 $0.180 / $0.590 DeepInfra
Meta-Llama-3.1-70B-Instruct-Turbo $0.100 / $0.280 $0.880 / $0.880 DeepInfra
Llama-3.3-70B-Instruct-Turbo $0.130 / $0.390 $0.880 / $0.880 DeepInfra
Qwen3 Next 80B A3B Instruct $0.140 / $1.40 $0.150 / $1.50 DeepInfra
Qwen3 Next 80B A3B Thinking $0.140 / $1.40 $0.150 / $1.50 DeepInfra
Llama 4 Maverick $0.150 / $0.600 $0.270 / $0.850 DeepInfra
DeepSeek-V3.1 $0.270 / $1.00 $0.600 / $1.70 DeepInfra
Qwen3 235B A22B Thinking 2507 $0.300 / $2.90 $0.650 / $3.00 DeepInfra
DeepSeek V3 $0.380 / $0.890 $1.25 / $1.25 DeepInfra
Qwen3-Coder-480b-A35b-Instruct $0.400 / $1.60 $2.00 / $2.00 DeepInfra
Mixtral-8x7B-Instruct-V0.1 $0.400 / $0.400 $0.600 / $0.600 DeepInfra
Kimi K2 $0.500 / $2.00 $1.00 / $3.00 DeepInfra
DeepSeek R1 $0.700 / $2.40 $3.00 / $7.00 DeepInfra

Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.