H100 SXM vs H100 PCIe

H100 SXM has 1.7× the bandwidth.

The NVIDIA H100 SXM (Hopper, 2022) and the NVIDIA H100 PCIe (Hopper, 2022) share the same Hopper silicon generation, so this comes down to configuration and price rather than architecture. Memory is effectively a wash at 80 vs 80 GB, so capacity does not decide this one. On memory bandwidth, the number that governs LLM serving speed, the H100 SXM leads at 3.35 vs 2 TB/s. Raw compute favors the H100 SXM by 1.3x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding.

Price is where it settles: the H100 PCIe rents from $1.60/hr against $1.90/hr for the H100 SXM, a 1.2x gap that the performance numbers above do justify for the right workload. Per gigabyte of VRAM, the H100 PCIe is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

NVIDIA · Hopper · 2022

NVIDIA H100 SXM

The workhorse of the AI boom. Still the most widely available serious training and inference GPU.

From $1.90/hr at Hyperstack

NVIDIA · Hopper · 2022

NVIDIA H100 PCIe

Hopper in a standard PCIe slot: ~75% of SXM compute at half the power, often meaningfully cheaper.

From $1.60/hr at Hyperstack

Head to head

■ H100 SXM   ■ H100 PCIe, bars share one scale across the whole catalog.

Memory 80 GB 80 GB
Memory bandwidth 3.35 TB/s 2 TB/s
FP16 dense 990 TFLOPS 756 TFLOPS
FP8 dense 1,979 TFLOPS 1,513 TFLOPS
Power (TDP) 700 W 350 W
H100 SXMH100 PCIe
Memory80 GB HBM380 GB HBM2e
Bandwidth3.35 TB/s2 TB/s
FP16 dense990 TF756 TF
FP8 dense1,979 TF1,513 TF
TDP700 W350 W
InterconnectNVLink 4 · 900 GB/sPCIe Gen5 · 128 GB/s (NVLink bridge optional)
Cheapest rental $1.90/hr $1.60/hr
$/hr per GB VRAM $2.4¢ $2.0¢

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.