Nemotron 3 Ultra API pricing
6 providers serve Nemotron 3 Ultra. Nemotron 3 Ultra is an open-weight model (OpenMDW-1.1) with 550B total parameters, 55B active per token (mixture-of-experts), up to 1049K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-10-08 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| W&B Inference | $0.500 | $2.15output floor | $0.100 | 262K | Use |
| DeepInfra | $0.500 | $2.20 | $0.100 | 262K | Use |
| OpenRouter | $0.500 | $2.20 | $0.100 | 262K | Use |
| Together AI | $0.600 | $3.60 | $0.200 | 512K | Use |
| Baseten | $0.600 | $2.40 | $0.120 | 203K | Use |
| Nebius | $1.00 | $3.00 | - | 1049K | Use |
The cheapest way to run Nemotron 3 Ultra
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $99.50/month, at W&B Inference. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, a busy self-hosted deployment can work out cheaper (~$1.51/1M output on 4× MI300X). See the math below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: Nemotron 3 Ultra self-hosting economics
Nemotron 3 Ultra is open-weight (OpenMDW-1.1), so the API price above competes with the GPU-hour market. 550B total parameters (MoE, ~55B active per token) needs roughly 757 GB VRAM at FP8 or 379 GB at INT4, KV-cache headroom included.
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| Nemotron 3 Nano | $0.050 | $0.200 | 4 |
| Nemotron 3.5 Lightning | $0.060 | $0.160 | 3 |
| Nemotron 3 Super | $0.080 | $0.400 | 3 |
| Nemotron 3.5 Content Safety | $0.200 | $0.200 | 1 |
FAQ
What is the cheapest API for Nemotron 3 Ultra?
As of 2026-10-08, the lowest input price for Nemotron 3 Ultra, $0.500 per 1M tokens, is the same list price at W&B Inference, DeepInfra and OpenRouter. The lowest output price is W&B Inference at $2.15 per 1M output tokens. The cheapest provider for a real workload depends on its input/output mix. The next-lowest input price is Together AI at $0.600, a 20% difference.
How much VRAM do you need to self-host Nemotron 3 Ultra?
Nemotron 3 Ultra has 550B parameters, so plan for roughly 757 GB of VRAM at FP8 or 379 GB at INT4/AWQ, KV-cache headroom included. That exceeds a single GPU. A typical node is 4× MI300X.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON