Will it fit?
The complete VRAM bill for serving any open model, weights, KV cache, and runtime overhead, and every GPU that clears it. Numbers come from each model's published config, not vibes.
The bill
GPUs that clear it
Assumes 90% of VRAM usable. Multi-GPU counts are for tensor-parallel serving (vLLM/SGLang/TGI). Prices: cheapest tracked on-demand, 2026-09.
| GPU | VRAM | GPUs needed | Est. $/hr | Cheapest at |
|---|
How this is computed. Weights = params × bytes/param × 1.08 (embeddings, norms, buffers). KV cache = per-token KV bytes from the model's config × context × sequences. Overhead = 1.5 GB CUDA/runtime + activation workspace. MoE models load all expert weights; sliding-window models (Gemma) need less KV than shown. Full derivation in the VRAM math.