NVIDIA · Grace Hopper · 2023
NVIDIA GH200
A Hopper GPU welded to a Grace CPU: 576 GB of unified fast memory for models that spill past VRAM.
96 GB
HBM3 + 480GB LPDDR5X
4 TB/s
Mem bandwidth
990 TF
FP16 dense
1,979 TF
FP8 dense
700 W
TDP
$1.35 /hr
From · Vast.ai
The NVLink-C2C link gives the GPU 900 GB/s access to 480 GB of CPU LPDDR5X, an order of magnitude faster than PCIe offload. Interesting for giant-model single-node inference and KV-cache offload.
Where to rent a GH200
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: GH200 pricing.
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Grace Hopper (2023) |
|---|---|
| Memory | 96 GB HBM3 + 480GB LPDDR5X, 4 TB/s |
| Interconnect | NVLink-C2C · 900 GB/s CPU↔GPU |
| Form factor | Superchip |
| Partitioning | MIG, up to 7 isolated instances |
What fits on one GH200
Weights + ~2 GB runtime overhead against 86 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Llama 3.3 70B | 70.6B | 8-bit |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | FP16 |
| Qwen2.5 32B | 32.8B | FP16 |
| Qwen2.5 Coder 32B | 32.8B | FP16 |
| Qwen2.5 72B | 72.7B | 8-bit |
| QwQ 32B (reasoning) | 32.8B | FP16 |
| Mistral 7B | 7.25B | FP16 |
| Mixtral 8x7B | 46.7B | 8-bit |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | FP16 |
| Phi-4 14B | 14.7B | FP16 |
| gpt-oss-20b | 20.9B | FP16 |
| gpt-oss-120b | 116.8B | 4-bit |
Best for
- CPU-offload inference of huge models
- Memory-elastic serving
- Single-node big-model experiments