Model
Llama 3.2 3B (3.2B) Llama 3.1 8B (8B) IBM Granite 4.1 8B (8.8B) Gemma 3 12B (12B) Phi-4 (14.7B) GPT-OSS 20B (21B, 3.6B active) Mistral Small 3.1 (24B) Gemma 3 27B (27B) Qwen3.6 27B (27.8B) Qwen3.5 27B (27.8B) Nemotron 3 Nano (30B, 3.5B active) IBM Granite 4.1 30B (30B) Qwen3 30B A3B (30.5B, 3.3B active) GLM-4.7 Flash (31B, 3B active) Qwen3 32B (32.8B) DeepSeek R1 Distill Qwen 32B (32.8B) Qwen3.6 35B A3B (36B, 3B active) Qwen3.5 35B A3B (36B, 3B active) Mixtral 8x7B (46.7B, 12.9B active) Llama 3.3 70B (70B) DeepSeek R1 Distill Llama 70B (70B) GLM-4.5 Air (106B, 12B active) Llama 4 Scout (109B, 17B active) GPT-OSS 120B (117B, 5.1B active) Leanstral 1.5 (119B, 6.5B active) Mistral Small 4 (119.4B) Nemotron 3 Super (120B, 12B active) Mistral Medium 3.5 (128B) MiniMax M2.7 (229B, 10B active) MiniMax M2.5 (229B, 10B active) MiniMax M2.1 (229B, 10B active) MiniMax M2 (229B, 10B active) Qwen3 235B A22B (235B, 22B active) DeepSeek V4 Flash (284B, 13B active) GLM-4.5 (355B, 32B active) GLM-4.6 (357B, 32B active) GLM-4.7 (358B, 32B active) Qwen3.5 397B A17B (397B, 17B active) Llama 4 Maverick (400B, 17B active) Llama 3.1 405B (405B) MiniMax M3 (428B, 23B active) Qwen3 Coder (480B, 35B active) Nemotron 3 Ultra (550B, 55B active) DeepSeek R1 (685B, 37B active) DeepSeek V3 (685B, 37B active) DeepSeek V3.2 (685B, 37B active) GLM-5.2 (744B, 40B active) GLM-5.1 (744B, 40B active) GLM-5 (744B, 40B active) Kimi K2 Thinking (1058B, 32B active) Kimi K2 (1058B, 32B active) Kimi K2.7 Code (1059B, 32B active) Kimi K2.6 (1059B, 32B active) Kimi K2.5 (1059B, 32B active) DeepSeek V4 Pro (1600B, 49B active)
Quantization
FP8 (~1.1 GB/B + headroom)
INT4 / AWQ (~0.55 GB/B + headroom)
GPU
A6000 · 48 GB · $0.287/hr (Vast.ai, marketplace)
6000 Ada · 48 GB · $0.388/hr (Vast.ai, marketplace)
L40S · 48 GB · $0.467/hr (Vast.ai, marketplace)
RTX PRO 6000 · 96 GB · $0.500/hr (RunPod, secure)
A100 SXM 80GB · 80 GB · $0.667/hr (Vast.ai, marketplace)
H100 SXM · 80 GB · $1.34/hr (Vast.ai, marketplace)
H200 · 141 GB · $2.00/hr (Verda, spot)
B200 · 180 GB · $3.06/hr (Verda, spot)
V100 · 32 GB · $0.022/hr (Vast.ai, marketplace)
A100 SXM 40GB · 40 GB · $0.388/hr (Vast.ai, marketplace)
B300 · 268 GB · $3.75/hr (Verda, spot)
RTX 3090 · 24 GB · $0.068/hr (Vast.ai, marketplace)
RTX 5080 · 16 GB · $0.102/hr (Vast.ai, marketplace)
RTX 4090 · 24 GB · $0.134/hr (Vast.ai, marketplace)
RTX 5090 · 32 GB · $0.321/hr (Vast.ai, marketplace)
L40 · 48 GB · $0.402/hr (Vast.ai, marketplace)
A100 PCIe · 80 GB · $0.442/hr (Vast.ai, marketplace)
H200 NVL · 141 GB · $0.500/hr (RunPod, community)
H100 NVL · 94 GB · $1.47/hr (Vast.ai, marketplace)
H100 PCIe · 80 GB · $1.80/hr (Vast.ai, marketplace)
RTX A2000 · 6 GB · $0.120/hr (RunPod, community)
RTX A5000 · 24 GB · $0.160/hr (RunPod, community)
RTX A4000 · 16 GB · $0.170/hr (RunPod, community)
RTX 4000 SFF Ada Generation · 20 GB · $0.180/hr (RunPod, community)
RTX 4070 Ti · 12 GB · $0.190/hr (RunPod, community)
RTX A4500 · 20 GB · $0.190/hr (RunPod, community)
RTX 4000 Ada Generation · 20 GB · $0.200/hr (RunPod, community)
RTX 2000 Ada Generation · 16 GB · $0.240/hr (RunPod, secure)
RTX 4080 · 16 GB · $0.270/hr (RunPod, community)
RTX PRO 4500 Blackwell · 32 GB · $0.340/hr (RunPod, community)
A40 · 48 GB · $0.350/hr (RunPod, community)
L4 · 24 GB · $0.440/hr (RunPod, community)
RTX 5000 Ada Generation · 32 GB · $0.490/hr (RunPod, community)
MI300X · 192 GB · $0.500/hr (RunPod, community)
RTX PRO 4000 Blackwell · 24 GB · $0.500/hr (RunPod, community)
RTX PRO 4500 Blackwell Server Edition · 32 GB · $0.500/hr (RunPod, community)
RTX PRO 5000 Blackwell · 48 GB · $0.820/hr (RunPod, community)
GB300 · 288 GB · $4.31/hr (Verda, spot)
RTX 3070 · 8 GB · $0.130/hr (RunPod, community)
RTX 3080 · 10 GB · $0.170/hr (RunPod, community)
RTX 3080 Ti · 12 GB · $0.180/hr (RunPod, community)
A10 · 24 GB · $0.241/hr (Vast.ai, marketplace)
Aggregate throughput (tok/s, batched)
Defaulted from active params. Replace with your own benchmark.
Utilization (% of rented hours doing useful work)
Tokens generated per month (millions)