H200 vs GH200
H200 has 1.5× the memory; H200 has 1.2× the bandwidth; GH200 rents for 1.7× less.
The NVIDIA H200 (Hopper, 2024) and the NVIDIA GH200 (Grace Hopper, 2023) are close contemporaries from adjacent generations, so the details decide. On capacity, the H200 carries 1.5x the memory (141 vs 96 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the H200 leads at 4.8 vs 4 TB/s.
Price is where it settles: the GH200 rents from $1.35/hr against $2.30/hr for the H200, a 1.7x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the GH200 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.
NVIDIA · Hopper · 2024
NVIDIA H200
An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.
From $2.30/hr at Nebius
NVIDIA · Grace Hopper · 2023
NVIDIA GH200
A Hopper GPU welded to a Grace CPU: 576 GB of unified fast memory for models that spill past VRAM.
From $1.35/hr at Vast.ai
Head to head
■ H200 ■ GH200, bars share one scale across the whole catalog.
| H200 | GH200 | |
|---|---|---|
| Memory | 141 GB HBM3e | 96 GB HBM3 + 480GB LPDDR5X |
| Bandwidth | 4.8 TB/s | 4 TB/s |
| FP16 dense | 990 TF | 990 TF |
| FP8 dense | 1,979 TF | 1,979 TF |
| TDP | 700 W | 700 W |
| Interconnect | NVLink 4 · 900 GB/s | NVLink-C2C · 900 GB/s CPU↔GPU |
| Cheapest rental | $2.30/hr | $1.35/hr |
| $/hr per GB VRAM | $1.6¢ | $1.4¢ |
What one can run that the other can't
Only on H200
- Mixtral 8x22B (4-bit)
Only on GH200
Nothing, H200 runs everything GH200 does.
Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.