The worked example
A team sizing Hunyuan MT2 30B-A3B for a translation pipeline does the memory math before requesting a single quote. The name says 30B total, A3B active, which reads like a small model. It isn’t, not for VRAM purposes.
Weights at FP8 run roughly 1.1 GB per billion parameters, so 30B total comes to 33 GB. Add the usual 25% headroom for KV cache and you land at 41.25 GB. That number is the whole story: it clears a 40 GB card. An A100 SXM 40GB is out. You need a 48 GB-class card at minimum, whether L40S, A6000, or RTX 6000 Ada, or you jump straight to an A100 80GB and stop worrying about it.
Drop to INT4 and the math changes fast. At 0.55 GB per billion params, 30B comes to 16.5 GB, and 25% headroom brings it to 20.6 GB. That fits an L4 or an RTX 4090 with room to spare. Same model, same active-param count, two completely different rental tiers depending on quantization choice. That gap, not the model’s headline size, is what decides your GPU bill.
Where the smaller two sizes land
The 7B dense variant is unremarkable by comparison. Its FP8 footprint runs 7.7 GB, or about 9.6 GB with headroom. That sits comfortably on an L4 with plenty of VRAM left for batching multiple concurrent requests, which matters more for a 7B chat or translation workload than raw capacity does.
The 1.8B dense variant barely registers. Its FP8 footprint lands under 2.5 GB. At that size, memory stops being the constraint entirely. Every card in the catalog, down to the smallest workstation SKUs like the RTX A2000, holds the weights with room left over. The question shifts from “what fits” to “what throughput do I need per dollar,” which is a different guide.
The trap in the middle tier
Hunyuan MT2 doesn’t have a live pricing row on /models yet, so there’s no way to check its per-token API rate directly. That absence is itself informative: until it lands, anyone planning around it is pricing an analog, not the model itself. The closest published stand-ins by architecture are Qwen3 30B A3B and Nemotron 3 Nano 30B A3B, both dense-total-30-ish-billion, both active around 3-3.5B. If MT2’s 30B-A3B follows that naming convention, its active count sits in the same range. That’s an inference from naming, not a confirmed spec. Treat it as a planning estimate, not a fact to build a capacity commitment on.
The trap engineers hit is assuming a 3B active count means 3B-scale hosting cost. It doesn’t. Every layer’s weights sit in VRAM regardless of how many experts fire per token, because the router can send any token to any expert on any forward pass. Total parameters pay the memory bill. That’s why 30B-A3B needs the same card tier as a 27-32B dense model at FP8, not the tier a genuinely 3B dense model would need. Anyone who quotes a 30B-A3B deployment against 3B-dense pricing has under-provisioned before the first request lands.
Once Hunyuan MT2 does get priced across providers, expect the spread to be wide. Established MoE models with a similarly lopsided active-to-total ratio already show this. GPT-OSS 120B, at 117B total against 5.1B active, currently spreads over 20x between cheapest and priciest provider on /models. Sparse models are cheap to serve at scale but expensive to serve badly, and that gap shows up directly in quoted rates. Don’t lock into the first MT2 provider you find without checking that page once it populates.
Where the discount is
If you’re running the INT4 build on a consumer card, the community and marketplace tier matters more than the model choice does. RTX 4090 rentals currently run 70%+ cheaper on marketplace or community listings than secure on-demand, as of this writing, per /gpus/rtx-4090. For a 20.6 GB footprint with no multi-tenant isolation requirement, that’s real money left on the table if you default to secure tier out of habit. It’s not a fit for anything you can’t tolerate losing mid-run. For a batch translation job with checkpointing, it’s the better default.
The decision rule
Compute your target MT2 size’s footprint with the 1.1 GB (FP8) or 0.55 GB (INT4) per-billion-param rule plus 25% headroom, before you look at any provider quote. If the number clears 40 GB, you’re buying a 48 GB+ card no matter which of the three sizes you picked, because that threshold is set by total parameters, not by the “A3B” label. If it lands under 24 GB, price the INT4 build against community RTX 4090 rates first: at this model’s memory footprint, the discount tier is doing more for your budget than the model choice is.