Tencent's Hy4 just got priced: where it lands against the Hunyuan lineup

Published 2026-09-16
Hy4 flagship bracket priced daily on llmhosting.ai

No row, no price

Tencent shipped Hy4 as the flagship sitting above its Hunyuan mt2 stack: a 30B-A3B mixture, a 7B dense model, and a 1.8B model underneath it. Check /models today and none of the four show up. That’s not an oversight on our end. The catalog tracks 3266 live price rows across providers, refreshed daily, and Hunyuan simply hasn’t been listed by a provider yet. If someone hands you a per-token rate for Hy4 right now, they’re quoting a rumor or an internal estimate, not a real listing.

That absence is the actual finding, not a gap to apologize for. It changes which question is worth answering. Instead of “what does Hy4 cost,” the useful question is “what should Hy4 cost, given the shape Tencent chose, and how do you avoid getting burned by the first quote once a provider does list it.”

The tiering pattern isn’t new

One flagship plus a ladder of smaller variants is the standard MoE playbook, and we track plenty of examples that follow it. GLM-4.7 runs 358B total against 32B active, while GLM-4.7 Flash drops to 31B total and 3B active, a roughly 10x cut in active compute for a model clearly meant to sit under the flagship in the same family. Kimi runs the same logic at bigger scale. Kimi K2.7 Code carries 1059B total params but only 32B active, a 3% active ratio that’s typical of frontier-scale MoE releases right now.

Hunyuan’s mt2 lineup (30B-A3B, 7B, 1.8B) already shows this pattern in miniature: the 30B-A3B tier is a sparse mixture with a small active slice, while the 7B and 1.8B are presumably dense and pay full compute per token. If Hy4 follows the family logic instead of breaking from it, expect it to land as a large sparse mixture with an active fraction well below its total, not a dense scale-up of the 7B model.

Sizing the flagship you can’t see yet

MiniMax M3428B total / 23B act~470GB FP8 Hy4 (unpriced)total params = ?active params = ? DeepSeek V4 Pro1600B total / 49B act~3x the VRAM if light bracket:single H200 maybe ok if heavy bracket:single-node OOMsunder load Rule: provision forthe heavier bracket,not the lighter one if mid-flagship if top-tier VRAM tracks total params; latency tracks active params — they move independently provision for the heavy bracket, not the light one

We don’t have Tencent’s real param counts for Hy4, and inventing a number would be worse than admitting we don’t have one. What we can do is bracket it against comparable tracked flagships. If Hy4 sits in the mid-flagship range, something like MiniMax M3 (428B total, 23B active) is a reasonable analog for VRAM planning: at FP8, that’s roughly 470GB before KV headroom, which pushes you toward multi-card builds on H200 rather than a single accelerator. If Hy4 instead scales toward the top of the field like DeepSeek V4 Pro (1600B total, 49B active), the memory bill triples even though the active-param compute bill barely moves, because total params drive VRAM and active params drive per-token latency largely independently.

That gap between the two brackets is the actual risk. Picture a team that provisions a single H100 node assuming Hy4 lands in the MiniMax-M3 weight class, then finds it’s closer to a 1T+ total-param design once the real card lands. It has under-provisioned by a wide enough margin that the node OOMs on the first concurrent multi-session load test, not on the first single request. That failure mode shows up during a demo, not during planning, which is the worst time to discover it.

What happens when the row finally lands

Qwen3 235B A22Bdozens of providers20x+ price spread Hy4, first week1-2 providersspread: worse, not better H200 NVL rentalssecure vs communitydiffer 70%+ first quote = ceiling,not the floor worse dispersionon new listings Rule: quote at least 3 providers before committing capacity 20x spread on a settled model — expect worse on day one

When a provider does list Hy4, don’t trust the first number you see. New flagship listings are exactly where cross-provider pricing is most inconsistent, because early adopters price against internal cost estimates instead of settled market rates. As of this writing, a mature, heavily-tracked model like Qwen3 235B A22B still shows a provider spread over 20x on the same weights, and that’s a model that’s been live for a long stretch with dozens of hosts competing on it. A model with one or two providers in its first week will show worse dispersion, not better.

The same logic applies on the hardware side if you’re weighing self-hosting instead of an API. Secure-tier and community-tier rentals on capable cards like H200 NVL routinely differ by more than 70%, so whatever secure-tier quote you get for running Hy4’s weight class isn’t the floor, it’s the ceiling.

The rule for today

If you need Hy4 pricing for a decision this week, don’t build a budget on any number you’ve seen quoted outside our tables, because none of those numbers are backed by a live provider row yet. Bracket your VRAM and compute plan using the nearest tracked analog to whichever total-param class Tencent eventually confirms, provision for the heavier bracket rather than the lighter one, and check /models daily until a real Hunyuan row appears. Once it does, run it against at least three providers before you commit capacity. The spread on a brand-new listing will be worse than anything you’re used to seeing on a settled model.

Put the numbers to work. Run your own workload through the breakeven calculator with today's floors, then price hardware: rent at Vast.ai → · rent at RunPod → Rental links are referral links. They never affect our rankings, which are ordered by price alone.

← All guides