RunPod
GPU cloud with per-second billing across Secure Cloud (T3/T4 datacenters) and cheaper Community Cloud, plus serverless GPU endpoints with autoscaling.
From frontier labs to peer-to-peer GPU marketplaces. Each profile shows live prices and where the provider actually wins. Price tables across the site also cover additional providers that have no profile yet.
GPU cloud with per-second billing across Secure Cloud (T3/T4 datacenters) and cheaper Community Cloud, plus serverless GPU endpoints with autoscaling.
Peer-to-peer GPU marketplace: the cheapest raw $/hr anywhere.
AI-first GPU cloud known for simple flat pricing, fast multi-node clusters and first access to new NVIDIA hardware.
GPU Droplets bring Paperspace's GPU lineup into DigitalOcean's developer-friendly cloud: H100s and L40S with predictable pricing, plus the broader DO ecosystem (storage, networking, managed databases).
Global cloud with 32 datacenter regions offering fractional and full GPUs (GH200, A100, L40S, A16).
Combines a serverless LLM API (DeepSeek, Llama, Qwen at aggressive per-token prices) with GPU instances and template deployments, one of the few providers covering both sides of the hosting equation.
US-based GPU cloud renting A100/H100-class VMs by the hour with straightforward pricing and pre-built NVIDIA-partner images.
Now an enterprise cluster provider: dedicated H100/H200/B200/GB200 clusters as an NVIDIA Cloud Partner, priced by quote only.
Marketplace of independent hosts offering deep-discount GPUs (explicitly no affiliate program to keep margins thin).
Hyperscale-class GPU cloud built on Kubernetes, favored for large training clusters (H100/H200/GB200 at scale) with InfiniBand networking.
AI cloud from the former Yandex team with H100/H200/B200 clusters in Europe and the US, plus Nebius AI Studio for per-token open-weight inference.
NexGen Cloud's self-service GPU platform running on 100% renewable energy, with H100/A100 VMs, NVLink options and minute-level billing in European and North American datacenters.
German dedicated-server stalwart offering RTX 4000/6000-class GPU dedicated servers at unbeatable monthly flat rates.
European cloud giant with budget GPU instances (V100, L4, L40S, H100) and strong data-sovereignty guarantees.
High-performance open-weight inference (Llama, DeepSeek, Qwen) on a custom stack, plus fine-tuning and GPU clusters.
Python-native serverless GPUs: decorate a function, get autoscaling H100s with sub-second cold starts.
The GPT and o-series API.
Claude models via first-party API.
Gemini models with the most generous free tier of any frontier lab and very cheap Flash-class models.
One API key for 340+ models across every major provider, with automatic failover and pass-through pricing.
Speed-focused open-model inference with FireAttention serving stack, function-calling models and image generation.
Custom LPU silicon serving open models at hundreds of tokens/second: the speed king for latency-sensitive apps, with simple per-token pricing.
Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options.
European frontier lab with strong small/medium models (and EU data processing), first-party API plus open weights you can self-host commercially.
Frontier-class reasoning and coding models at a fraction of Western lab prices, with off-peak discounts and open weights for self-hosting.
Grok models via API with live X data access and competitive frontier pricing tiers.
Wafer-scale silicon serving open models at extreme speed (2000+ tokens/sec on some models).
RDU-based inference cloud serving large open models (405B-class) at high speed with per-token pricing.
Claude, Llama, Mistral, Nova and more inside AWS with IAM, VPC and compliance controls.
Gemini plus 100+ models in Google Cloud's ML platform, with enterprise governance, grounding and tuning pipelines.
OpenAI models with Azure enterprise controls plus a broad third-party model catalog.
Run thousands of community models (LLMs, image, audio, video) behind one API with per-second GPU billing: the fastest path from 'found a model on GitHub' to production endpoint.
Production inference platform with Truss packaging, optimized serving engines and enterprise-grade autoscaling.
Inference Endpoints for dedicated deployments plus Inference Providers routing to partner clouds, all integrated with the Hub where the models already live.