H200 vs Gaudi 3

H200 has 1.3× the bandwidth.

The NVIDIA H200 (Hopper, 2024) and the Intel Gaudi 3 (Gaudi, 2024) come from different vendors and different design goals, but they get cross-shopped for a reason. Memory is effectively a wash at 141 vs 128 GB, so capacity does not decide this one. On memory bandwidth, the number that governs LLM serving speed, the H200 leads at 4.8 vs 3.7 TB/s.

Price is where it settles: the H200 rents from $2.30/hr against $2.30/hr for the Gaudi 3, close enough that price should not decide. Per gigabyte of VRAM, the H200 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. The full case for and against Intel's part is in Gaudi 3: the accelerator nobody is fighting over. Everything below is the underlying data.

NVIDIA · Hopper · 2024

NVIDIA H200

An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.

From $2.30/hr at Nebius

Intel · Gaudi · 2024

Intel Gaudi 3

Intel's price-performance play, with Ethernet-native scale-out instead of proprietary interconnect.

From $2.30/hr at FluidStack

Head to head

■ H200   ■ Gaudi 3, bars share one scale across the whole catalog.

Memory 141 GB 128 GB
Memory bandwidth 4.8 TB/s 3.7 TB/s
FP16 dense 990 TFLOPS 918 TFLOPS
FP8 dense 1,979 TFLOPS 1,835 TFLOPS
Power (TDP) 700 W 900 W
H200Gaudi 3
Memory141 GB HBM3e128 GB HBM2e
Bandwidth4.8 TB/s3.7 TB/s
FP16 dense990 TF918 TF
FP8 dense1,979 TF1,835 TF
TDP700 W900 W
InterconnectNVLink 4 · 900 GB/s24×200GbE RoCE on-chip
Cheapest rental $2.30/hr $2.30/hr
$/hr per GB VRAM $1.6¢ $1.8¢

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.