H100 SXM vs B200

B200 has 2.4× the memory; B200 has 2.4× the bandwidth; H100 SXM rents for 2.0× less.

The NVIDIA H100 SXM (Hopper, 2022) and the NVIDIA B200 (Blackwell, 2024) are 2 years and a silicon generation apart, which makes the price gap the heart of the story. On capacity, the B200 carries 2.4x the memory (192 vs 80 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the B200 leads at 8 vs 3.35 TB/s. Raw compute favors the B200 by 2.3x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding.

Price is where it settles: the H100 SXM rents from $1.90/hr against $3.75/hr for the B200, a 2.0x gap that the performance numbers above do justify for the right workload. Per gigabyte of VRAM, the B200 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

NVIDIA · Hopper · 2022

NVIDIA H100 SXM

The workhorse of the AI boom. Still the most widely available serious training and inference GPU.

From $1.90/hr at Hyperstack

NVIDIA · Blackwell · 2024

NVIDIA B200

NVIDIA's Blackwell flagship: 192 GB of HBM3e and roughly double Hopper's throughput per chip.

From $3.75/hr at Packet.ai

Head to head

■ H100 SXM   ■ B200, bars share one scale across the whole catalog.

Memory 80 GB 192 GB
Memory bandwidth 3.35 TB/s 8 TB/s
FP16 dense 990 TFLOPS 2,250 TFLOPS
FP8 dense 1,979 TFLOPS 4,500 TFLOPS
Power (TDP) 700 W 1,000 W
H100 SXMB200
Memory80 GB HBM3192 GB HBM3e
Bandwidth3.35 TB/s8 TB/s
FP16 dense990 TF2,250 TF
FP8 dense1,979 TF4,500 TF
TDP700 W1000 W
InterconnectNVLink 4 · 900 GB/sNVLink 5 · 1.8 TB/s
Cheapest rental $1.90/hr $3.75/hr
$/hr per GB VRAM $2.4¢ $2.0¢

What one can run that the other can't

Only on H100 SXM

Nothing, B200 runs everything H100 SXM does.

Only on B200

  • Mixtral 8x22B (8-bit)
  • gpt-oss-120b (8-bit)

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.