NVIDIA · Grace Hopper · 2023

NVIDIA GH200

A Hopper GPU welded to a Grace CPU: 576 GB of unified fast memory for models that spill past VRAM.

96 GB
HBM3 + 480GB LPDDR5X
4 TB/s
Mem bandwidth
990 TF
FP16 dense
1,979 TF
FP8 dense
700 W
TDP
$1.35 /hr
From · Vast.ai

The NVLink-C2C link gives the GPU 900 GB/s access to 480 GB of CPU LPDDR5X, an order of magnitude faster than PCIe offload. Interesting for giant-model single-node inference and KV-cache offload.

Where to rent a GH200

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: GH200 pricing.

Provider $/GPU-hr vs cheapest Notes
Vast.ai $1.35 cheapest Limited listings
Lambda $1.49 1.1× Promotional single-node pricing

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 96 GB
Memory bandwidth 4 TB/s
FP16 dense 990 TFLOPS
FP8 dense 1,979 TFLOPS
Power (TDP) 700 W
ArchitectureGrace Hopper (2023)
Memory96 GB HBM3 + 480GB LPDDR5X, 4 TB/s
InterconnectNVLink-C2C · 900 GB/s CPU↔GPU
Form factorSuperchip
PartitioningMIG, up to 7 isolated instances

What fits on one GH200

Weights + ~2 GB runtime overhead against 86 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Llama 3.3 70B 70.6B 8-bit
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B FP16
Qwen2.5 32B 32.8B FP16
Qwen2.5 Coder 32B 32.8B FP16
Qwen2.5 72B 72.7B 8-bit
QwQ 32B (reasoning) 32.8B FP16
Mistral 7B 7.25B FP16
Mixtral 8x7B 46.7B 8-bit
Gemma 2 9B 9.24B FP16
Gemma 2 27B 27.2B FP16
Phi-4 14B 14.7B FP16
gpt-oss-20b 20.9B FP16
gpt-oss-120b 116.8B 4-bit

Best for

Compare

H100 SXM vs GH200

The workhorse of the AI boom. Still the most widely available serious training and inference GPU.

H200 vs GH200

An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.