NVIDIA · Hopper · 2022

NVIDIA H100 PCIe

Hopper in a standard PCIe slot: ~75% of SXM compute at half the power, often meaningfully cheaper.

80 GB
HBM2e
2 TB/s
Mem bandwidth
756 TF
FP16 dense
1,513 TF
FP8 dense
350 W
TDP
$1.60 /hr
From · Hyperstack

Lower clocks and HBM2e instead of HBM3 mean about 60% of the SXM's memory bandwidth. Fine for single-GPU work; the weaker interconnect hurts multi-GPU training.

Decision guide: A100 vs H100 for inference →

Where to rent a H100 PCIe

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: H100 PCIe pricing.

Provider $/GPU-hr vs cheapest Notes
Hyperstack $1.60 cheapest neocloud
FluidStack $1.85 1.2× neocloud
Vast.ai $2.56 1.6× Verified-listing rate, refreshed daily
RunPod $2.89 1.8× neocloud
OVHcloud $2.99 1.9× hyperscaler

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 80 GB
Memory bandwidth 2 TB/s
FP16 dense 756 TFLOPS
FP8 dense 1,513 TFLOPS
Power (TDP) 350 W
ArchitectureHopper (2022)
Memory80 GB HBM2e, 2 TB/s
InterconnectPCIe Gen5 · 128 GB/s (NVLink bridge optional)
Form factorPCIe
PartitioningMIG, up to 7 isolated instances

What fits on one H100 PCIe

Weights + ~2 GB runtime overhead against 72 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Llama 3.3 70B 70.6B 4-bit
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B FP16
Qwen2.5 32B 32.8B 8-bit
Qwen2.5 Coder 32B 32.8B 8-bit
Qwen2.5 72B 72.7B 4-bit
QwQ 32B (reasoning) 32.8B 8-bit
Mistral 7B 7.25B FP16
Mixtral 8x7B 46.7B 8-bit
Gemma 2 9B 9.24B FP16
Gemma 2 27B 27.2B FP16
Phi-4 14B 14.7B FP16
gpt-oss-20b 20.9B FP16

Best for

Compare

H100 SXM vs H100 PCIe

The workhorse of the AI boom. Still the most widely available serious training and inference GPU.