GPUs · Clouds · Pricing · 2026-09
Find the cheapest way to run any AI model.
Virtualized tracks 21 accelerators and 123 live prices across 24 clouds, then turns them into straight answers: which GPU fits your model, what it costs per token, and whether to rent or own.
Cheapest on-demand today
full GPU pricing →| GPU | VRAM | From | Where |
|---|---|---|---|
| H100 SXM | 80 GB | $1.90/hr | Hyperstack |
| H200 | 141 GB | $2.30/hr | Nebius |
| B200 | 192 GB | $3.75/hr | Packet.ai |
| MI300X | 192 GB | $1.85/hr | Vast.ai |
| A100 80GB | 80 GB | $0.87/hr | Vast.ai |
| RTX 4090 | 24 GB | $0.34/hr | RunPod |
On-demand list prices per GPU-hour, last reviewed 2026-09. Full tables: GPU pricing · all 24 clouds · methodology.
Tools that answer the real questions
Calculator
Will this model fit on that GPU?
Pick a model, precision, and context length. Get the VRAM bill, weights, KV cache, overhead, and every GPU that clears it, with the cheapest rental for each.
Estimator
What does self-hosting actually cost per million tokens?
Throughput-based $/Mtok for any model on any GPU at real utilization, the number to hold against API pricing before you rent anything.
Calculator
Buy the GPU or keep renting it?
Purchase price + power + utilization → effective $/hr, payback period, and break-even against every rental rate we track.
The matchups
all comparisons →Understand the layer
all guides →The A100 in 2026: the best value in AI inference
Six years old and unbeatable on bandwidth-per-dollar, the arithmetic behind the market's favorite value GPU.
Cutting inference cost per token
The seven levers that move $/token, ranked by impact, starting with the latency-vs-throughput split.
The VRAM math every AI engineer should know
Weights, KV cache, activations, the three numbers that decide whether a model fits, worked by hand.
Why your tokens/sec is 10× lower than the benchmark
Continuous batching, KV cache paging, and the memory-bandwidth wall, explained with pictures.
The AI compute stack, from sand to serverless
Eight layers between a transistor and your API call, and who takes margin at each one.