GLM, GLM Flash, or GLM-5.3 Flash: pricing Zhipu's three live tiers

Published 2026-09-28
GLM tiers pricing gap priced daily on llmhosting.ai

Two tabs open, one deadline. GLM-5.2 in one, GLM-4.7 Flash in the other, and a ticket that says “wire up the agent loop by end of day.” Somewhere in a planning doc someone wrote “GLM-5.3 Flash” as the target model. That third tab won’t load, because there’s no live pricing row for it yet.

The active-param gap is the whole story

GLM-5.2744B total40B active GLM-4.7 Flash31B total3B active ~13x active-param gap rate ratio~=13xcompute-honest rate ratio<< 13xFlash subsidized,default to Flash check live rates check live rates Active params = compute bill = your per-token rate. Total params only set the hosting memory bill. 13x is the test, not the label

GLM-5.2 runs at 744B total parameters with 40B active per token. GLM-4.7 Flash runs at 31B total with 3B active. That’s roughly a 13x gap in active params, and active params are what you pay for on every forward pass, not total params. Total params set the memory bill for hosting the weights; active params set the compute bill for generating each token. On an API, the compute bill is the one that shows up as your per-token rate.

So run a live test right now instead of trusting anyone’s memory of last quarter’s pricing. Pull today’s per-token rate off GLM-5.2 and off GLM-4.7 Flash, and check whether the ratio between them tracks the 13x active-param ratio. If it does, the pricing is compute-honest and you’re choosing on capability, not on a subsidy. If the gap is much narrower than 13x, Zhipu is pricing Flash as a loss-leading adoption tier, and Flash becomes the correct default for almost everything that doesn’t need flagship reasoning depth.

Why the third tab is empty

GLM-5.3 Flash doesn’t have a pricing row because, as of this writing, Zhipu hasn’t shipped a Flash variant onto the 5.x reasoning line. The naming pattern that exists is a mid-tier: GLM-4.5 Air sits at 106B total, 12B active, between the full 4.5 line and the 3B-active Flash tier that later shipped as 4.7. If Zhipu eventually ships a 5.x Flash, the Air-to-Flash compression (12B active down to 3B) is the only precedent you have for guessing where it lands, and it’s one data point from one generation. That’s not a pattern, it’s an anecdote. Don’t build a budget line around a model that hasn’t cleared a pricing row on /models. Check back there before you commit a roadmap item to it.

Routing chat vs agent traffic on what’s live today

Chat traffic is mostly single-shot: one prompt, one completion, low context carryover. Active-param compute dominates the cost, and Flash’s 3B active tier is cheap enough that quality is the only reason to pay more. If your evals show GLM-4.7 Flash clearing your accuracy bar on chat-shaped tasks, route there by default.

Agent traffic is different. Multi-step tool calls re-feed context every turn, and a shallow model that has to retry a malformed tool call or a missed constraint burns tokens fast enough to erase Flash’s rate advantage inside a single session. So skip the sticker price and measure steps-to-completion on a matched sample of your real tasks. Run twenty representative agent runs through both tiers, count turns and retries, and multiply by each tier’s current rate off the two model pages above. Flagship’s premium buys fewer retries; whether that’s worth it depends on your retry rate today, not on the label “flagship.”

The self-hosting crossover, and where it breaks

GLM-5.2 FP8~818GB weights+25% KV >1TB multi-cardclusterH200 / H100 SXM GLM-4.7 FlashINT4 ~17GB+headroom ~21GB single cardL4 / RTX 4090room to spare 2x H100 SXMfits ONE streamat ~511GB concurrentsessions stackKV -> mid-runOOM single-stream math add concurrency Flash's small footprint, not its cheap label, is why it self-hosts on a rentable card. Same split, opposite shopping list

If you’re weighing self-hosting instead of the API, the same active-vs-total split flips which resource you’re shopping for. GLM-5.2’s weights at FP8 run about 744B times 1.1GB per billion params, roughly 818GB, plus 25% KV headroom pushes you past 1TB. That clears a single card by a wide margin and turns into a multi-card cluster decision priced off H200 or H100 SXM rows, not a single-box one.

GLM-4.7 Flash at INT4 is a different animal: 31B times 0.55GB is about 17GB, plus headroom lands near 21GB, which fits on a single L4 or RTX 4090 with room to spare. That’s the crossover most people get backwards. They assume the “Flash” label means it’s not worth self-hosting, when it’s actually the tier cheap enough to run on a card you can rent by the hour without a cluster contract.

Where this bites: teams that try to squeeze GLM-5.2 onto two H100 SXM cards at INT4, figuring 409GB plus 25% headroom (about 511GB) fits two cards’ combined capacity with room to spare. The math holds for one request at a time. The moment concurrent sessions climb past a handful, KV cache for each in-flight context stacks on top of that headroom estimate, and the box OOMs mid-run, not at startup, which is the worst time to find out your capacity math was single-stream.

The rule

Default to GLM-4.7 Flash for chat and for any agent task where a twenty-run sample shows retries staying low. Escalate to GLM-5.2 only when the matched-sample test shows Flash’s retry rate eating its rate advantage. Do not write GLM-5.3 Flash into a budget or a roadmap until it clears a live row on /models; check /trends periodically if you’re waiting on it, and price your cluster capacity for concurrent sessions, not for the single-stream number the memory formula gives you.

Put the numbers to work. Run your own workload through the breakeven calculator with today's floors, then price hardware: rent at Vast.ai → · rent at RunPod → Rental links are referral links. They never affect our rankings, which are ordered by price alone.

← All guides