H200 cloud pricing, August 2026

Live H200 quotes (141 GB VRAM) across the providers we track, cheapest first. Marketplace prices are real asks from the Vast.ai order book; secure/on-demand tiers carry datacenter SLAs (what the tiers mean).

Quotes refreshed 2026-08-23
Price floor: $2.00/hr (spot) at Verda.
Provider Tier $/hr per GPU Median $/hr Live offers
Verda spot $2.00 - -
Vast.ai marketplace $3.29 $3.95 28 Rent
RunPod community $3.59 - - Rent
Verda on-demand $4.00 - -
RunPod secure $4.59 - - Rent
Need 8-128 of these? Get a cluster quote →, or hand us the workload for a ranked deployment plan, free while in beta.

One H200 or two H100 SXMs?

For models needing more than 80 GB but at most 141 GB, a single H200 holds the whole model on one card: no tensor-parallel split, no inter-GPU communication overhead. A pair of H100 SXMs brings more raw compute but pays that overhead. Today's quotes, same provider and tier on both sides (the 2× column doubles the per-GPU price):

ProviderTier1× H200 $/hr2× H100 SXM $/hrCheaper
Verda spot $2.00 $3.25 H20038.5% less
Vast.ai marketplace $3.29 $2.67 2× H100 SXM18.9% less
RunPod community $3.59 $5.38 H20033.3% less
Verda on-demand $4.00 $6.50 H20038.5% less
RunPod secure $4.59 $6.58 H20030.2% less

From today's quotes: one H200 is cheaper in 4 of 5 comparable provider/tier pairs. Recomputed on every daily refresh.

What fits on a single H200

Models from our tracked open-weight set that fit in 141 GB, with KV-cache headroom included. Bigger models need multi-GPU nodes (math in the calculator).

ModelParamsFits atVRAM neededCheapest API $/1M out
GLM-4.7 Flash 31B FP8 43 GB compare →
Qwen3.6 35B A3B 36B FP8 50 GB compare →
Qwen3.6 27B 27.8B FP8 39 GB compare →
Qwen3.5 35B A3B 36B FP8 50 GB compare →
Qwen3.5 27B 27.8B FP8 39 GB compare →
Qwen3 32B 32.8B FP8 46 GB compare →
Qwen3 30B A3B 30.5B FP8 42 GB compare →
Mistral Small 3.1 24B FP8 33 GB
Nemotron 3 Nano 30B FP8 42 GB compare →
IBM Granite 4.1 30B 30B FP8 42 GB
IBM Granite 4.1 8B 8.8B FP8 13 GB compare →
GPT-OSS 20B 21B FP8 29 GB compare →
Llama 3.3 70B 70B FP8 97 GB compare →
Llama 3.1 8B 8B FP8 11 GB compare →
Llama 3.2 3B 3.2B FP8 5 GB compare →
Gemma 3 27B 27B FP8 38 GB compare →
Gemma 3 12B 12B FP8 17 GB compare →
Phi-4 14.7B FP8 21 GB compare →
Mixtral 8x7B 46.7B FP8 65 GB compare →
DeepSeek R1 Distill Llama 70B 70B FP8 97 GB compare →
DeepSeek R1 Distill Qwen 32B 32.8B FP8 46 GB compare →
GLM-4.5 Air 106B INT4 73 GB compare →
Mistral Medium 3.5 128B INT4 88 GB compare →
Mistral Small 4 119.4B INT4 83 GB
Leanstral 1.5 119B INT4 82 GB
Nemotron 3 Super 120B INT4 83 GB compare →
GPT-OSS 120B 117B INT4 81 GB compare →
Llama 4 Scout 109B INT4 75 GB compare →

Can't buy a H200? The desk-side alternatives

H200 cards only ship in servers, but the "own your inference box" itch has real answers: run the economics in the calculator, then look at what this audience actually buys:

GMKtec EVO-X2 (128GB unified)

Ryzen AI Max+ 395 · 128GB unified memory

The $/GB winner: runs 70B-class and quantized big-MoE (V4 Flash tier) that no consumer GPU fits

NVIDIA RTX 5090 (32GB)

32GB GDDR7 · ~1.8TB/s bandwidth

Bandwidth-per-dollar king for the A3B-MoE and 27-32B dense workhorses

As an Amazon Associate we earn from qualifying purchases · hardware links may be referral links · picks are editorial, not paid · disclosure

FAQ

How much does it cost to rent a H200?

As of 2026-08-23, H200 rentals start at $2.00/hr (spot tier at Verda). Datacenter capacity with SLAs typically costs more than community or marketplace hardware.

What can you run on a H200?

With 141 GB of VRAM, a single H200 fits models up to roughly 100B parameters at FP8 or 201B at INT4 quantization, including KV-cache headroom.

Cite or embed today's floor

Live badge for a README. It updates itself with every daily refresh:

H200 rental floor badge

[![H200 rental floor](https://llmhosting.ai/badge/h200.svg)](https://llmhosting.ai/gpus/h200)

Citation: H200 rental floor $2.00/hr (llmhosting.ai, 2026-08-23). Free to reuse with a link back.

Daily H200 price history →