RTX 4000 SFF Ada Generation cloud pricing, August 2026
Live RTX 4000 SFF Ada Generation quotes (20 GB VRAM) across the providers we track, cheapest first. Marketplace prices are real asks from the Vast.ai order book; secure/on-demand tiers carry datacenter SLAs (what the tiers mean).
Quotes refreshed 2026-08-23| Provider | Tier | $/hr per GPU | Median $/hr | Live offers | |
|---|---|---|---|---|---|
| RunPod | community | $0.180 | - | - | Rent |
| RunPod | secure | $0.440 | - | - | Rent |
What fits on a single RTX 4000 SFF Ada Generation
Models from our tracked open-weight set that fit in 20 GB, with KV-cache headroom included. Bigger models need multi-GPU nodes (math in the calculator).
| Model | Params | Fits at | VRAM needed | Cheapest API $/1M out |
|---|---|---|---|---|
| IBM Granite 4.1 8B | 8.8B | FP8 | 13 GB | compare → |
| Llama 3.1 8B | 8B | FP8 | 11 GB | compare → |
| Llama 3.2 3B | 3.2B | FP8 | 5 GB | compare → |
| Gemma 3 12B | 12B | FP8 | 17 GB | compare → |
| Qwen3.6 27B | 27.8B | INT4 | 20 GB | compare → |
| Qwen3.5 27B | 27.8B | INT4 | 20 GB | compare → |
| Mistral Small 3.1 | 24B | INT4 | 17 GB | |
| GPT-OSS 20B | 21B | INT4 | 15 GB | compare → |
| Gemma 3 27B | 27B | INT4 | 19 GB | compare → |
| Phi-4 | 14.7B | INT4 | 11 GB | compare → |
FAQ
How much does it cost to rent a RTX 4000 SFF Ada Generation?
As of 2026-08-23, RTX 4000 SFF Ada Generation rentals start at $0.180/hr (community tier at RunPod). Datacenter capacity with SLAs typically costs more than community or marketplace hardware.
What can you run on a RTX 4000 SFF Ada Generation?
With 20 GB of VRAM, a single RTX 4000 SFF Ada Generation fits models up to roughly 14B parameters at FP8 or 28B at INT4 quantization, including KV-cache headroom.
Cite or embed today's floor
Live badge for a README. It updates itself with every daily refresh:
[](https://llmhosting.ai/gpus/rtx-4000-sff-ada-generation)
Citation: RTX 4000 SFF Ada Generation rental floor $0.180/hr (llmhosting.ai, 2026-08-23). Free to reuse with a link back.