NVIDIA · Hopper · 2022

NVIDIA H100 SXM

The workhorse of the AI boom. Still the most widely available serious training and inference GPU.

80 GB
HBM3
3.35 TB/s
Mem bandwidth
990 TF
FP16 dense
1,979 TF
FP8 dense
700 W
TDP
$1.90 /hr
From · Hyperstack

The GPU that trained most of the current generation of frontier models. Transformer Engine FP8, 3.35 TB/s of HBM3, NVLink for multi-GPU scale-out. Rental prices have fallen dramatically since 2023 as supply caught up.

Decision guide: A100 vs H100 for inference →

Where to rent a H100 SXM

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: H100 SXM pricing.

Provider $/GPU-hr vs cheapest Notes
Hyperstack $1.90 cheapest neocloud
Nebius $1.95 1.0× neocloud
Genesis $1.99 1.0× neocloud
FluidStack $2.10 1.1× neocloud
Vast.ai $2.14 1.1× Verified-listing rate, refreshed daily; interruptible lower
DataCrunch $2.19 1.2× neocloud
TensorDock $2.25 1.2× marketplace
Together $2.39 1.3× neocloud
Crusoe $2.45 1.3× neocloud
Lambda $2.49 1.3× neocloud
Cudo $2.49 1.3× marketplace
RunPod $2.69 1.4× Secure Cloud; Community from ~$1.99
Massed $2.79 1.5× neocloud
DigitalOcean $2.99 1.6× neocloud
Jarvislabs $2.99 1.6× neocloud
CoreWeave $3.35 1.8× neocloud
Modal $3.95 2.1× Serverless, per-second billing
OCI $4.00 2.1× BM.GPU.H100.8 ÷ 8
AWS $4.90 2.6× p5.48xlarge ÷ 8, post-2025 price cut
Azure $5.70 3.0× ND H100 v5 ÷ 8
GCP $5.90 3.1× A3 ÷ 8

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 80 GB
Memory bandwidth 3.35 TB/s
FP16 dense 990 TFLOPS
FP8 dense 1,979 TFLOPS
Power (TDP) 700 W
ArchitectureHopper (2022)
Memory80 GB HBM3, 3.35 TB/s
InterconnectNVLink 4 · 900 GB/s
Form factorSXM
PartitioningMIG, up to 7 isolated instances

What fits on one H100 SXM

Weights + ~2 GB runtime overhead against 72 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Llama 3.3 70B 70.6B 4-bit
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B FP16
Qwen2.5 32B 32.8B 8-bit
Qwen2.5 Coder 32B 32.8B 8-bit
Qwen2.5 72B 72.7B 4-bit
QwQ 32B (reasoning) 32.8B 8-bit
Mistral 7B 7.25B FP16
Mixtral 8x7B 46.7B 8-bit
Gemma 2 9B 9.24B FP16
Gemma 2 27B 27.2B FP16
Phi-4 14B 14.7B FP16
gpt-oss-20b 20.9B FP16

Best for

Compare

H100 SXM vs H200

An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.

H100 SXM vs A100 80GB

The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.

H100 SXM vs B200

NVIDIA's Blackwell flagship: 192 GB of HBM3e and roughly double Hopper's throughput per chip.

H100 SXM vs MI300X

AMD's answer to Hopper: 192 GB on one GPU, a 70B model in FP16 fits with room to spare.

H100 SXM vs H100 PCIe

Hopper in a standard PCIe slot: ~75% of SXM compute at half the power, often meaningfully cheaper.

RTX 5090 vs H100 SXM

The fastest thing you can put in a desktop: 32 GB of GDDR7 at 1.8 TB/s.

H100 SXM vs Gaudi 3

Intel's price-performance play, with Ethernet-native scale-out instead of proprietary interconnect.

H100 SXM vs GH200

A Hopper GPU welded to a Grace CPU: 576 GB of unified fast memory for models that spill past VRAM.

TPU v6e vs H100 SXM

Trillium: ~4.7× v5e compute per chip. Google's mainstream training silicon.