The setup that loses money on day one
A team picks Granite 4.2 8B because it’s small, IBM-licensed, and cheap on paper. They rent an RTX PRO 6000 to self-host it, reasoning that an 8B model plus a flagship workstation card is obviously overkill in the good direction. Three weeks later the card is running at single-digit utilization most of the day, because their actual traffic is a few million tokens a month, not a few million an hour. They’re now paying for 730 hours of a card every month to do work that would fit into a few dozen hours of continuous inference. The API would have been cheaper, and it wasn’t close.
That’s the failure mode this math exists to catch: small model, wrong GPU tier, wrong volume assumption, all three compounding.
The absence is the finding
Granite 4.2 8B doesn’t have a live pricing row on /models yet. What’s tracked today is IBM Granite 4.1 8B, a dense 8.8B-parameter model, same architecture class, same weight class, the closest real analog you can price against right now. Since 4.1 is dense (8.8B total, 8.8B active, no MoE sparsity to complicate the memory math), it’s a clean stand-in for sizing the hardware side of this decision even before 4.2’s API rate shows up. Check the Granite 4.1 8B page for the current per-million rate before you run this against your own volume. That number is the other half of the equation, and it moves.
Why the RTX PRO 6000 is the wrong default
Run the standard rule of thumb: roughly 1.1 GB of VRAM per billion parameters at FP8, plus 25% headroom for KV cache. For an 8.8B dense model that’s about 9.7 GB of weights, call it 12-13 GB fully loaded with a reasonable context window. That fits on an L4 with room to spare. It fits on an A6000 with room to spare too. It does not need the memory bandwidth or the price tier of an RTX PRO 6000, which is built for models running into the hundreds of billions of total parameters, not a single dense 8.8B checkpoint.
Renting the bigger card for a model this small isn’t caution, it’s a scheduling penalty. You pay the card’s hourly rate whether it’s processing a token every second or a token every ten minutes. The API bills per token processed. That asymmetry is the entire crux of this decision, and it means GPU choice is secondary to volume. Get the volume assumption wrong and no GPU choice fixes it.
The variable that decides this
The breakeven isn’t a fixed number, it’s a ratio: the GPU’s monthly rental floor divided by the model’s per-million-token rate, converted into a token count. Pull the current hourly rate on /gpus/l4, multiply by 730 hours to get your monthly hardware floor, then check the live rate on the Granite 4.1 8B page to see how many million tokens of API usage that floor buys you. Below that volume, the API wins outright. Above it, self-hosting starts paying for itself, provided the GPU is actually busy and not idling between requests.
That second condition is where teams get burned even after doing the volume math correctly. A GPU rental that clears breakeven on paper only clears it in practice if your traffic is sustained rather than bursty. A card sitting at 15% utilization for three weeks and then getting hammered for two days doesn’t average out to a good deal; it averages out to a bad deal with occasional good days. The calculator runs this with today’s live rate on both sides, but it can’t know your actual request pattern. You have to supply that.
What the market tells you about picking providers, even for models this small
Provider spread on comparably sized dense models is wide enough that picking the wrong host erases any self-hosting advantage before you’ve rented anything. Llama 3.1 8B, a similar dense 8B checkpoint, shows more than 20x spread between its cheapest and priciest listed provider on its page as of this writing. If Granite 4.2 lands with anything like that spread once it’s live, the provider you pick matters more than the self-host-versus-API decision itself. Check /models for the current row before assuming the sticker rate you saw once is still the floor.
What this can’t tell you
It can’t tell you IBM’s actual API rate for 4.2 before that row goes live. And it can’t tell you your own request pattern, whether your traffic is steady background load or spiky demo traffic that’ll leave a rented card idle most of the week. Both of those numbers are yours to pull, not ours to assume.
The decision rule
Before renting anything, compute your monthly token volume against the current Granite 4.1 8B rate and the current L4 hourly rate, not the RTX PRO 6000 rate, since the model doesn’t need that card. If your volume clears the breakeven and your traffic is sustained rather than bursty, rent the L4 or an equivalent small card and self-host. If either condition is soft, stay on the API until your volume or your consistency catches up, and rerun the check monthly on /trends since GPU rates and API rates don’t move together.