Rent a GPU, or pay per token?

The only question that matters is utilization. This calculator takes today's live GPU floors and API floors, your expected throughput, and tells you where the lines cross. Every field is editable. Our defaults are planning estimates, not benchmarks. Prefer a ready-made answer? The self-hosting cost table runs this math for every notable open model at once. Rather hand the whole decision off? Describe your workload and we'll send back a ranked API-vs-rent-vs-own plan, free while in beta.

GPU floors refreshed 2026-08-23

Your workload

The verdict

Keep the workload in-house

If the model fits and the machine will run for years, buying can beat both API tokens and rented GPU-hours. Our local-hardware calculator includes live purchase prices, electricity, usable AI memory, and a cash break-even against today's cloud floor.

Compare buy vs rent →

Method: GPU count = VRAM needed ÷ card VRAM (rounded up, tensor parallelism assumed free; it isn't quite). Self-host $/1M = (GPU count × $/hr) ÷ (throughput × 3600 × utilization) × 10⁶. API comparison uses today's cheapest listed output price. Input-token costs, egress, storage and engineering time are not included.