Describe your workload. Get a deployment plan.
One form, one reply: three ranked ways to run your workload. Use an API, rent GPUs, or own the box. Each option comes with monthly and annual cost, GPU count and VRAM, break-even volume, price timestamps, a confidence label, and why it ranked where it did. Built from the same price data this site refreshes daily.
Free while in beta. Prepared individually. A human reviews every request, checks every number, and replies by email.
Your workload
answer what you know; unknowns are part of the jobWhat a plan looks like
illustrative example, not your numbersExample workload: Llama-70B-class chat, ~150M output tokens/mo, interactive latency.
| # | Path | Monthly | Annual | Hardware | Break-even | Confidence | Why |
|---|---|---|---|---|---|---|---|
| 1 | Use an API | ~$95 | ~$1,140 | none | wins below ~1B tok/mo | high | 15× cheaper than renting at this volume; zero ops |
| 2 | Rent GPUs | ~$1,450 | ~$17,400 | 1× H100 80GB (INT4) | wins past ~1B tok/mo | medium | full data control; only pays off at higher volume or 24/7 batching |
| 3 | Own the box | ~$60 power | ~$720 + $9,500 upfront | 2× RTX 4090 | ~7 months vs renting, 24/7 | low | cheapest long-run if you'll operate it; tight VRAM fit, real ops burden |
Prefer self-serve? The breakeven calculator and the self-hosting cost table run this math interactively from today's live prices.