NVIDIA · Hopper · 2022
NVIDIA H100 SXM
The workhorse of the AI boom. Still the most widely available serious training and inference GPU.
The GPU that trained most of the current generation of frontier models. Transformer Engine FP8, 3.35 TB/s of HBM3, NVLink for multi-GPU scale-out. Rental prices have fallen dramatically since 2023 as supply caught up.
Decision guide: A100 vs H100 for inference →
Where to rent a H100 SXM
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: H100 SXM pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| Hyperstack | $1.90 | neocloud |
| Nebius | $1.95 | neocloud |
| Genesis | $1.99 | neocloud |
| FluidStack | $2.10 | neocloud |
| Vast.ai | $2.14 | Verified-listing rate, refreshed daily; interruptible lower |
| DataCrunch | $2.19 | neocloud |
| TensorDock | $2.25 | marketplace |
| Together | $2.39 | neocloud |
| Crusoe | $2.45 | neocloud |
| Lambda | $2.49 | neocloud |
| Cudo | $2.49 | marketplace |
| RunPod | $2.69 | Secure Cloud; Community from ~$1.99 |
| Massed | $2.79 | neocloud |
| DigitalOcean | $2.99 | neocloud |
| Jarvislabs | $2.99 | neocloud |
| CoreWeave | $3.35 | neocloud |
| Modal | $3.95 | Serverless, per-second billing |
| OCI | $4.00 | BM.GPU.H100.8 ÷ 8 |
| AWS | $4.90 | p5.48xlarge ÷ 8, post-2025 price cut |
| Azure | $5.70 | ND H100 v5 ÷ 8 |
| GCP | $5.90 | A3 ÷ 8 |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Hopper (2022) |
|---|---|
| Memory | 80 GB HBM3, 3.35 TB/s |
| Interconnect | NVLink 4 · 900 GB/s |
| Form factor | SXM |
| Partitioning | MIG, up to 7 isolated instances |
What fits on one H100 SXM
Weights + ~2 GB runtime overhead against 72 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Llama 3.3 70B | 70.6B | 4-bit |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | FP16 |
| Qwen2.5 32B | 32.8B | 8-bit |
| Qwen2.5 Coder 32B | 32.8B | 8-bit |
| Qwen2.5 72B | 72.7B | 4-bit |
| QwQ 32B (reasoning) | 32.8B | 8-bit |
| Mistral 7B | 7.25B | FP16 |
| Mixtral 8x7B | 46.7B | 8-bit |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | FP16 |
| Phi-4 14B | 14.7B | FP16 |
| gpt-oss-20b | 20.9B | FP16 |
Best for
- Training and fine-tuning at every scale
- Inference for models up to ~70B (single GPU, 4-bit)
- Multi-GPU clusters via NVLink/InfiniBand
Compare
H100 SXM vs H200
An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.
H100 SXM vs A100 80GB
The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.
H100 SXM vs B200
NVIDIA's Blackwell flagship: 192 GB of HBM3e and roughly double Hopper's throughput per chip.
H100 SXM vs MI300X
AMD's answer to Hopper: 192 GB on one GPU, a 70B model in FP16 fits with room to spare.
H100 SXM vs H100 PCIe
Hopper in a standard PCIe slot: ~75% of SXM compute at half the power, often meaningfully cheaper.
RTX 5090 vs H100 SXM
The fastest thing you can put in a desktop: 32 GB of GDDR7 at 1.8 TB/s.
H100 SXM vs Gaudi 3
Intel's price-performance play, with Ethernet-native scale-out instead of proprietary interconnect.
H100 SXM vs GH200
A Hopper GPU welded to a Grace CPU: 576 GB of unified fast memory for models that spill past VRAM.
TPU v6e vs H100 SXM
Trillium: ~4.7× v5e compute per chip. Google's mainstream training silicon.