GMKtec EVO-X2 (128GB unified)
The $/GB winner: runs 70B-class and quantized big-MoE (V4 Flash tier) that no consumer GPU fits
Live B300 quotes (288 GB VRAM) across the providers we track, cheapest first. Marketplace prices are real asks from the Vast.ai order book; secure/on-demand tiers carry datacenter SLAs (what the tiers mean).
Quotes refreshed 2026-08-23| Provider | Tier | $/hr per GPU | Median $/hr | Live offers | |
|---|---|---|---|---|---|
| Verda | spot | $3.75 | - | - | |
| RunPod | community | $6.94 | - | - | Rent |
| Verda | on-demand | $7.50 | - | - | |
| RunPod | secure | $7.89 | - | - | Rent |
Models from our tracked open-weight set that fit in 288 GB, with KV-cache headroom included. Bigger models need multi-GPU nodes (math in the calculator).
| Model | Params | Fits at | VRAM needed | Cheapest API $/1M out |
|---|---|---|---|---|
| GLM-4.7 Flash | 31B | FP8 | 43 GB | compare → |
| GLM-4.5 Air | 106B | FP8 | 146 GB | compare → |
| Qwen3.6 35B A3B | 36B | FP8 | 50 GB | compare → |
| Qwen3.6 27B | 27.8B | FP8 | 39 GB | compare → |
| Qwen3.5 35B A3B | 36B | FP8 | 50 GB | compare → |
| Qwen3.5 27B | 27.8B | FP8 | 39 GB | compare → |
| Qwen3 32B | 32.8B | FP8 | 46 GB | compare → |
| Qwen3 30B A3B | 30.5B | FP8 | 42 GB | compare → |
| Mistral Medium 3.5 | 128B | FP8 | 176 GB | compare → |
| Mistral Small 4 | 119.4B | FP8 | 165 GB | |
| Leanstral 1.5 | 119B | FP8 | 164 GB | |
| Mistral Small 3.1 | 24B | FP8 | 33 GB | |
| Nemotron 3 Super | 120B | FP8 | 165 GB | compare → |
| Nemotron 3 Nano | 30B | FP8 | 42 GB | compare → |
| IBM Granite 4.1 30B | 30B | FP8 | 42 GB | |
| IBM Granite 4.1 8B | 8.8B | FP8 | 13 GB | compare → |
| GPT-OSS 120B | 117B | FP8 | 161 GB | compare → |
| GPT-OSS 20B | 21B | FP8 | 29 GB | compare → |
| Llama 4 Scout | 109B | FP8 | 150 GB | compare → |
| Llama 3.3 70B | 70B | FP8 | 97 GB | compare → |
| Llama 3.1 8B | 8B | FP8 | 11 GB | compare → |
| Llama 3.2 3B | 3.2B | FP8 | 5 GB | compare → |
| Gemma 3 27B | 27B | FP8 | 38 GB | compare → |
| Gemma 3 12B | 12B | FP8 | 17 GB | compare → |
| Phi-4 | 14.7B | FP8 | 21 GB | compare → |
| Mixtral 8x7B | 46.7B | FP8 | 65 GB | compare → |
| DeepSeek R1 Distill Llama 70B | 70B | FP8 | 97 GB | compare → |
| DeepSeek R1 Distill Qwen 32B | 32.8B | FP8 | 46 GB | compare → |
| DeepSeek V4 Flash | 284B | INT4 | 196 GB | compare → |
| GLM-4.7 | 358B | INT4 | 247 GB | compare → |
| GLM-4.6 | 357B | INT4 | 246 GB | compare → |
| GLM-4.5 | 355B | INT4 | 245 GB | compare → |
| MiniMax M2.7 | 229B | INT4 | 158 GB | compare → |
| MiniMax M2.5 | 229B | INT4 | 158 GB | compare → |
| MiniMax M2.1 | 229B | INT4 | 158 GB | compare → |
| MiniMax M2 | 229B | INT4 | 158 GB | compare → |
| Qwen3.5 397B A17B | 397B | INT4 | 273 GB | compare → |
| Qwen3 235B A22B | 235B | INT4 | 162 GB | compare → |
| Llama 4 Maverick | 400B | INT4 | 276 GB | compare → |
| Llama 3.1 405B | 405B | INT4 | 279 GB | compare → |
B300 cards only ship in servers, but the "own your inference box" itch has real answers: run the economics in the calculator, then look at what this audience actually buys:
The $/GB winner: runs 70B-class and quantized big-MoE (V4 Flash tier) that no consumer GPU fits
Bandwidth-per-dollar king for the A3B-MoE and 27-32B dense workhorses
The startup standard: only sub-$10K single card that runs 70B+ comfortably
As an Amazon Associate we earn from qualifying purchases · hardware links may be referral links · picks are editorial, not paid · disclosure
As of 2026-08-23, B300 rentals start at $3.75/hr (spot tier at Verda). Datacenter capacity with SLAs typically costs more than community or marketplace hardware.
With 288 GB of VRAM, a single B300 fits models up to roughly 205B parameters at FP8 or 411B at INT4 quantization, including KV-cache headroom.
Live badge for a README. It updates itself with every daily refresh:
[](https://llmhosting.ai/gpus/b300)
Citation: B300 rental floor $3.75/hr (llmhosting.ai, 2026-08-23). Free to reuse with a link back.