Baseten pricing & review
Production inference platform with Truss packaging, optimized serving engines and enterprise-grade autoscaling. Strong for teams shipping custom models with SLAs.
Where Baseten wins
- Optimized model serving (TensorRT-LLM)
- Enterprise autoscaling and observability
Baseten model pricing
List prices refreshed 2026-10-08 ยท cheapest 25 shown| Model | Input $/1M | Output $/1M | Market floor in |
|---|---|---|---|
| GPT-OSS 120B | $0.100 | $0.500 | $0.030at W&B Inference |
| DeepSeek V4 Flash | $0.130 | $0.260 | $0.030at OpenRouter |
| GLM 5.3 Flash | $0.150 | $0.500 | $0.110at Sail |
| DeepSeek V4.1 Flash | $0.300 | $1.20 | $0.044at OpenRouter |
| MiniMax M2.5 | $0.300 | $1.20 | $0.270at OpenRouter |
| DeepSeek-V3.1 | $0.500 | $1.50 | $0.250at DeepInfra |
| Inkling Small | $0.500 | $1.20 | $0.450at DeepInfra |
| GLM-4.7 | $0.600 | $2.20 | $0.400at GMI Cloud |
| Kimi K2.5 | $0.600 | $3.00 | $0.450at OpenRouter |
| Kimi K2 Thinking | $0.600 | $2.50 | $0.600at AWS Bedrock |
| Kimi K2 | $0.600 | $2.50 | $0.500at DeepInfra |
| GLM-4.6 | $0.600 | $2.20 | $0.430at OpenRouter |
| Nemotron 3 Ultra | $0.600 | $2.40 | $0.500at W&B Inference |
| DeepSeek V3 | $0.770 | $0.770 | $0.200at Hyperbolic |
| Kimi K2.6 | $0.950 | $4.00 | $0.650at W&B Inference |
| Kimi K2.7 Code | $0.950 | $4.00 | $0.671at OpenRouter |
| GLM-5 | $0.950 | $3.15 | $0.600at OpenRouter |
| Inkling | $1.00 | $4.05 | $0.950at DeepInfra |
| GLM-5.2 | $1.40 | $4.40 | $0.152at OpenRouter |
| GLM 5.3 | $1.40 | $4.40 | $0.070at OpenRouter |
| DeepSeek V4 Pro | $1.74 | $3.48 | $0.209at OpenRouter |
| GLM-5p2-Fast | $2.10 | $6.60 | $2.10at Fireworks AI |
| Kimi K3 | $3.00 | $15.00 | $0.790at OpenRouter |