About
Virtualized is an independent reference for the virtualized AI compute layer: the GPUs, the clouds that rent them, and the arithmetic that connects a model to a hardware bill.
Why this exists
AI compute is the strangest market on the internet. The same physical GPU rents for prices that differ by 3-4× depending on the vendor. Capability questions with exact answers, will a 70B model fit on this card?, are mostly answered by folklore. And the vocabulary of the layer (MIG, microVMs, paged KV caches) is documented for specialists only. This site closes those three gaps: a spec database, a price index, and calculators plus explainers that show their work.
Price methodology
Indicative on-demand list prices per GPU-hour, normalized to a single GPU (multi-GPU instance prices divided by GPU count), marketplace rates refreshed daily from public provider APIs (25th percentile of verified single-GPU on-demand listings), list prices last reviewed September 2026. Spot, interruptible, reserved, and negotiated pricing is typically 30-70% lower. Always confirm on the provider's pricing page before committing.
Marketplace and community-cloud rates (Vast.ai, RunPod) refresh automatically every day from public provider APIs, with sanity guards that reject implausible values rather than publish them. Fixed list prices (hyperscalers, neoclouds without public rate APIs) get a full manual review monthly. Every price on the site carries the cadence it was collected under.
- We track 123 prices across 24 providers and 21 accelerators, last full review 2026-09.
- Multi-GPU instance prices are divided by GPU count to a per-GPU-hour figure.
- Marketplace prices (Vast.ai) are the median of active verified listings at review time, not the outlier floor.
- We list on-demand only. Reserved, spot, and negotiated pricing is routinely 30-70% lower and changes the ranking, treat our table as the ceiling you should never pay more than.
Calculator methodology
The VRAM calculator computes weights (params × bytes × 1.08), KV cache from each model's published architecture (layers, KV heads, head dim), and a fixed runtime overhead, against 90% of each GPU's VRAM. The inference cost estimator models decode throughput as the minimum of a bandwidth bound (50% MBU) and a compute bound (40% MFU), deliberately conservative, real stacks with good batching sometimes beat it. Full derivations are in the VRAM math and the serving explainer.
Independence & disclosure
No provider pays for placement, ordering is always by price or spec, and outbound links carry no affiliate tags. One relationship to know about: Virtualized is operated by the team behind Charg, a GPU redeployment and hosting business. Charg does not appear in the price index or rankings, and where we link to it in editorial content, we say so inline. Every number on this site stays checkable against public sources regardless.
Corrections
Spotted a stale price or a wrong spec? The entire dataset is versioned in public, corrections are welcome and usually shipped within a day.