The RAM price shock broke 2025's self-hosting math

Published 2026-07-24
DDR5 price shock priced daily on llmhosting.ai

The rig that used to make sense

DDR5 spot prices are up several hundred percent from their 2024 lows as of this writing. That single fact quietly gutted a piece of conventional wisdom that circulated all through 2024 and 2025: buy a workstation board, stuff it with 512GB or more of system RAM, drop in one consumer GPU, and run huge mixture-of-experts models via CPU offload for a fraction of API or rental cost. The GPU handled the active experts. RAM held everything else, cold, waiting to be paged in. It worked because DDR5 was the cheap side of that equation. It no longer is.

A hobbyist who priced out a 512GB build in 2024 and only pulled the trigger this year found the RAM line item alone now costs several times more than the GPU sitting next to it. That inversion, RAM as the expensive component instead of the throwaway one, is the whole story.

Why total params, not active params, drove the RAM bill

Kimi K21059B total / 32B active~580GB cold experts DeepSeek V4 Pro1600B total / 49B active~880GB cold experts GLM-5.2744B total / 40B active~410GB cold experts System RAMholds ALL of it(cold, paged in) Consumer GPUholds onlyactive path pays RAM bill pays RAM bill pays RAM bill pages in active experts RAM cost scales with TOTAL params, not active params — that's the whole shock. cold experts still need a home

Total params pay the memory bill regardless of how few of them fire per token. Kimi K2 carries 1059B total params against 32B active. DeepSeek V4 Pro is 1600B total against 49B active. GLM-5.2 is 744B total against 40B active. Using the rough INT4 estimate of 0.55GB per billion params, the cold experts alone need on the order of 580GB, 880GB, and 410GB respectively, before you add the active path or any KV headroom. That capacity has to live somewhere. CPU offload setups, the kind vLLM and Ollama both support in some form, put it in system RAM specifically because RAM used to be the cheapest byte available. Today it isn’t a discount tier anymore. It sits closer to HBM than it used to, just still slower.

The consumer GPU ceiling didn’t move to compensate

Meanwhile the other side of the DIY rig, the GPU, stalled. Check the current RTX 4090 floor: 24GB. The RTX 5090 tops out around 32GB. Cards with real headroom above that, the RTX PRO 6000 or an A6000, sit well past the price ceiling most people call “enthusiast.” Run the FP8 math (roughly 1.1GB per billion active params) on Kimi K2’s 32B active path and you get about 35GB before KV cache, already over a 32GB card. Drop to INT4 and it’s roughly 18GB, which fits, but now you’ve quantized the compute path just to make room, and the 580GB of cold experts still needs a home. RAM was supposed to be that home at negligible marginal cost. That cost is no longer negligible.

Where the breakeven moved

Monthly utilization low → high 2024 breakevenlow utilizationstill favored owning 2026 breakevenneeds very highutilization to own Rent H200 / H100 SXMHBM holds fullactive path, no offload Dense models(Llama 3.3 70B, Qwen3 32B)barely affected RAM inflation shifts it renting wins wider range Shift hits large sparse MoE hardest; dense models don't need offload RAM at all. breakeven point, before vs after

This is the part that changes the buy-versus-rent decision, not just the parts list. The classic framing treats utilization as the only variable: high utilization favors owning hardware, low utilization favors renting by the hour. That framing assumed the owning side had a roughly fixed, low cost basis. RAM inflation broke that assumption specifically for large-MoE self-hosting, because the RAM requirement scales with total params, not active params, and total params are exactly where models like Kimi K2, DeepSeek V4 Pro, and GLM-5.2 are largest. A dense model like Llama 3.3 70B or Qwen3 32B never needed offload RAM at that scale in the first place, so this shock barely touches its economics. It hits the big sparse models hardest, which is exactly the class of model the “cheap home rig” advice was written for.

Feed the new RAM-adjusted build cost into the calculator against the live rate for an H200, which has enough HBM to hold a 40-50B active path plus KV cache without touching system RAM at all, and the monthly utilization threshold where owning wins climbs substantially. It hasn’t disappeared. It moved into a range most home builders and small teams won’t clear.

The part people still get backward

The instinct is to assume the consumer GPU itself is now the expensive leg of a DIY rig, since GPU prices get all the attention. It isn’t, right now. Marketplace and community tiers on cards like the RTX 4090 and RTX 5090 run at discounts of 70% or steeper against their secure on-demand tier. Renting that same card by the hour, at the marketplace rate, is frequently cheaper than the annualized cost of owning one once you price in a RAM-heavy chassis around it. The GPU was never the problem. The RAM was, and it moved fast enough that last year’s parts list is this year’s expensive mistake.

What to check before you build anything

If your target model’s total-param footprint at INT4 exceeds roughly 256GB, the CPU-offload rig is now competing against renting an H200 or H100 SXM node that holds the whole active path in HBM with no offload at all, and you should run both sides through the calculator with today’s DDR5-adjusted build cost before ordering hardware. Your regional RAM spot price and the current rental floor on the GPU list are the two numbers that settle it, and neither is printed here. Get both from the same week before you commit six months of amortization to a bill of materials that was priced for a different market.

Put the numbers to work. Run your own workload through the breakeven calculator with today's floors, then price hardware: rent at Vast.ai → · rent at RunPod → Rental links are referral links. They never affect our rankings, which are ordered by price alone.

← All guides