H100 SXM vs H200

H200 has 1.8× the memory; H200 has 1.4× the bandwidth; H100 SXM rents for 1.2× less.

The NVIDIA H100 SXM (Hopper, 2022) and the NVIDIA H200 (Hopper, 2024) share the same Hopper silicon generation, so this comes down to configuration and price rather than architecture. On capacity, the H200 carries 1.8x the memory (141 vs 80 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the H200 leads at 4.8 vs 3.35 TB/s.

Price is where it settles: the H100 SXM rents from $1.90/hr against $2.30/hr for the H200, a 1.2x gap that the performance numbers above do justify for the right workload. Per gigabyte of VRAM, the H200 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

NVIDIA · Hopper · 2022

NVIDIA H100 SXM

The workhorse of the AI boom. Still the most widely available serious training and inference GPU.

From $1.90/hr at Hyperstack

NVIDIA · Hopper · 2024

NVIDIA H200

An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.

From $2.30/hr at Nebius

Head to head

■ H100 SXM   ■ H200, bars share one scale across the whole catalog.

Memory 80 GB 141 GB
Memory bandwidth 3.35 TB/s 4.8 TB/s
FP16 dense 990 TFLOPS 990 TFLOPS
FP8 dense 1,979 TFLOPS 1,979 TFLOPS
Power (TDP) 700 W 700 W
H100 SXMH200
Memory80 GB HBM3141 GB HBM3e
Bandwidth3.35 TB/s4.8 TB/s
FP16 dense990 TF990 TF
FP8 dense1,979 TF1,979 TF
TDP700 W700 W
InterconnectNVLink 4 · 900 GB/sNVLink 4 · 900 GB/s
Cheapest rental $1.90/hr $2.30/hr
$/hr per GB VRAM $2.4¢ $1.6¢

What one can run that the other can't

Only on H100 SXM

Nothing, H200 runs everything H100 SXM does.

Only on H200

  • Mixtral 8x22B (4-bit)
  • gpt-oss-120b (4-bit)

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.