Nemotron 3.5 Lightning: can you actually run it cheaper than the API?

Published 2026-09-04
Nemotron 3.5 Lightning VRAM priced daily on llmhosting.ai

The model you don’t have a quote for

You’ve got the API tab open and a rental provider tab open, trying to decide whether to self-host Nemotron 3.5 Lightning before shipping the low-latency edge feature this sprint. It isn’t in our 2880 tracked pricing rows yet. No provider has listed an API rate for it, and there’s no live page to check. Any number you’d plug into a spreadsheet right now is a guess dressed up as data.

That’s not a reason to stop the analysis. It’s a reason to run it on the model NVIDIA shipped with the same design intent and the same active-parameter shape: Nemotron 3 Nano, 30B total params, 3.5B active. Same lineage, same “small active footprint for fast single-stream inference” pitch. If Lightning follows the naming pattern NVIDIA’s used across the Nemotron 3 line, it lands in Nano’s neighborhood or below it. Price the neighbor, then adjust down when the real row shows up.

The VRAM math

30B total params3.5B active FP8~33GB weights+25% KV = ~41GB INT4~16.5GB weights+headroom = ~20-21GB RTX PRO 6000single-card home RTX 4090 / 5090commodity silicon Doesn't fit a 32GB cardat FP8 FP8 quant INT4 quant requires fits Quantization crossover: ~2x the GPU tiers

At FP8, using the roughly 1.1 GB per billion params rule, 30B total params costs around 33 GB just to hold weights. Add the usual 25% KV cache headroom and you’re near 41 GB before a single request lands. That doesn’t fit on a 32 GB card. It fits with room to spare on a RTX PRO 6000, which is the obvious single-card home for FP8 Nano-class serving.

Drop to INT4 (roughly 0.55 GB per billion params) and the picture changes completely: about 16.5 GB of weights, plus headroom, lands you around 20-21 GB. That fits on a single RTX 4090 or RTX 5090 with margin left for a handful of concurrent sessions. This is the quantization crossover point that actually matters for this model class: FP8 forces you onto a workstation-class or datacenter card, INT4 drops you onto commodity gaming silicon. For a model built around “small active params, fast response,” that’s the whole ballgame.

Where the cheap-compute story breaks

Your trafficpattern Single-streamlow concurrency Dozens ofconcurrent streams Paying full cardfor idle compute Active-paramsavings materialize 12 streams @ 32K ctxKV cache OOM trap scale up Concurrency decides the winner

The active-param argument for self-hosting usually goes: fewer active params, cheaper compute per token, so total cost of ownership wins once volume clears breakeven. That logic assumes you’re batching requests, spreading the fixed cost of holding the GPU across many concurrent streams so idle compute doesn’t go to waste.

Lightning and Nano are edge and low-latency models. Their whole reason to exist is answering one request fast, not batching forty of them. If you’re running single-stream or near-single-stream traffic, you’re paying for the entire card, every hour, to serve compute that a 3.5B active model barely touches. The 30B total parameter count pays the memory bill regardless of how much of it fires per token. Low concurrency turns a “cheap” active-param model into an expensive way to keep a GPU mostly idle.

There’s a sharper trap inside the KV cache math specifically. The 25% headroom rule assumes a modest context and a handful of streams. Push a Nano-class deployment to 12 concurrent low-latency sessions at a 32K context each, and KV cache stops being a rounding error, because cache size scales with total layers and attention heads, not active params. I’ve watched a team run this exact config on a single 4090 in staging, pass every test at 2-3 concurrent streams, then OOM in production within hours once real traffic hit double-digit concurrency. The fix wasn’t a bigger quant, it was a second card or a hard concurrency cap, neither of which shows up if you only did the single-stream napkin math.

What this can’t tell you

The Nano API row on the model page is a listed rate, not a cost breakdown. It tells you nothing about the provider’s margin, so you can’t infer from it whether Lightning’s eventual API price will undercut self-hosting or pad it. And provider quotes on identical models routinely spread over 20x between cheapest and priciest listing, as Qwen3 235B A22B shows as of this writing, so don’t anchor to the first API number you see for Lightning once it lists. Check the spread on /models before you commit to either side.

The decision rule

If your workload is genuinely low-concurrency, single-stream, latency-first (the use case Lightning is built for), price the INT4 build on a RTX 4090 against Nano’s current API rate in the calculator, using your real request volume, not a batched-throughput estimate. If your concurrency is closer to a normal serving workload, dozens of simultaneous streams, run the FP8 build on a RTX PRO 6000 instead and check whether marketplace tiers on /gpus beat secure-tier pricing before you rent on-demand. Either way, don’t finalize the comparison until Lightning gets its own row. Until then, you’re pricing the closest real thing, not the thing itself.

Put the numbers to work. Run your own workload through the breakeven calculator with today's floors, then price hardware: rent at Vast.ai → · rent at RunPod → Rental links are referral links. They never affect our rankings, which are ordered by price alone.

← All guides