Will it fit?

The complete VRAM bill for serving any open model, weights, KV cache, and runtime overhead, and every GPU that clears it. Numbers come from each model's published config, not vibes.

The bill

GPUs that clear it

Assumes 90% of VRAM usable. Multi-GPU counts are for tensor-parallel serving (vLLM/SGLang/TGI). Prices: cheapest tracked on-demand, 2026-09.

GPUVRAMGPUs neededHeadroomEst. $/hrCheapest at

How this is computed. Weights = params × bytes/param × 1.08 (embeddings, norms, buffers). KV cache = per-token KV bytes from the model's config × context × sequences. Overhead = 1.5 GB CUDA/runtime + activation workspace. MoE models load all expert weights; sliding-window models (Gemma) need less KV than shown. Full derivation in the VRAM math.