Gemma 3 27B API pricing
7 providers serve Gemma 3 27B. Gemma 3 27B is an open-weight model (Gemma) with 27B total parameters, up to 262K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-08-23 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| Nebius | $0.060 | $0.200 | - | 128K | Use |
| OpenRouter marketplace quote | $0.080 | $0.450 | $0.040 | 262K | Use |
| DeepInfra | $0.090 | $0.160output floor | - | 131K | Use |
| Novita AI | $0.119 | $0.200 | - | 98K | Use |
| AWS Bedrock | $0.230 | $0.380 | - | 128K | Use |
| Scaleway | $0.250 | $0.500 | - | 40K | Use |
| Fireworks AI | $0.900 | $0.900 | - | 131K | Use |
The cheapest way to run Gemma 3 27B
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $10.20/month, at Nebius. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, the API floor remains below our self-hosting estimate of ~$0.239/1M. Details below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: Gemma 3 27B self-hosting economics
Gemma 3 27B is open-weight (Gemma), so the API price above competes with the GPU-hour market. 27B total parameters needs roughly 38 GB VRAM at FP8 or 19 GB at INT4, KV-cache headroom included.
| GPU that fits (FP8, single card) | VRAM | From $/hr | Cheapest at | |
|---|---|---|---|---|
| A6000 | 48 GB | $0.287marketplace | Vast.ai | Rent |
| A40 | 48 GB | $0.350community | RunPod | Rent |
| 6000 Ada | 48 GB | $0.388marketplace | Vast.ai | Rent |
| A100 SXM 40GB | 40 GB | $0.388marketplace | Vast.ai | Rent |
| L40 | 48 GB | $0.402marketplace | Vast.ai | Rent |
| A100 PCIe | 80 GB | $0.442marketplace | Vast.ai | Rent |
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| Gemma 3 12B | $0.050 | $0.100 | 5 |
| Gemma 4 26B A4B | $0.070 | $0.300 | 4 |
| Gemma 4 31B | $0.100 | $0.340 | 4 |
| Gemma-7b-It | $0.050 | $0.080 | 3 |
| Gemma 3 4B | $0.040 | $0.080 | 3 |
FAQ
What is the cheapest API for Gemma 3 27B?
As of 2026-08-23, the lowest input price for Gemma 3 27B is Nebius at $0.060 per 1M input tokens. The lowest output price is DeepInfra at $0.160 per 1M output tokens. The cheapest provider for a real workload depends on its input/output mix. The next-lowest input price is OpenRouter at $0.080, a 33% difference.
How much VRAM do you need to self-host Gemma 3 27B?
Gemma 3 27B has 27B parameters, so plan for roughly 38 GB of VRAM at FP8 or 19 GB at INT4/AWQ, KV-cache headroom included. A single A6000 (48 GB) fits it.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON