H200 vs Gaudi 3
H200 has 1.3× the bandwidth.
The NVIDIA H200 (Hopper, 2024) and the Intel Gaudi 3 (Gaudi, 2024) come from different vendors and different design goals, but they get cross-shopped for a reason. Memory is effectively a wash at 141 vs 128 GB, so capacity does not decide this one. On memory bandwidth, the number that governs LLM serving speed, the H200 leads at 4.8 vs 3.7 TB/s.
Price is where it settles: the H200 rents from $2.30/hr against $2.30/hr for the Gaudi 3, close enough that price should not decide. Per gigabyte of VRAM, the H200 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. The full case for and against Intel's part is in Gaudi 3: the accelerator nobody is fighting over. Everything below is the underlying data.
NVIDIA · Hopper · 2024
NVIDIA H200
An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.
From $2.30/hr at Nebius
Intel · Gaudi · 2024
Intel Gaudi 3
Intel's price-performance play, with Ethernet-native scale-out instead of proprietary interconnect.
From $2.30/hr at FluidStack
Head to head
■ H200 ■ Gaudi 3, bars share one scale across the whole catalog.
| H200 | Gaudi 3 | |
|---|---|---|
| Memory | 141 GB HBM3e | 128 GB HBM2e |
| Bandwidth | 4.8 TB/s | 3.7 TB/s |
| FP16 dense | 990 TF | 918 TF |
| FP8 dense | 1,979 TF | 1,835 TF |
| TDP | 700 W | 900 W |
| Interconnect | NVLink 4 · 900 GB/s | 24×200GbE RoCE on-chip |
| Cheapest rental | $2.30/hr | $2.30/hr |
| $/hr per GB VRAM | $1.6¢ | $1.8¢ |
Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.