NVIDIA · Hopper · 2024
NVIDIA H200
An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.
Same Hopper compute silicon as the H100 SXM, upgraded to 141 GB HBM3e at 4.8 TB/s. Because LLM inference is usually memory-bound, the H200 often delivers 1.4-1.9× H100 throughput on large models despite identical FLOPS.
Where to rent a H200
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: H200 pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| Nebius | $2.30 | neocloud |
| Hyperstack | $2.75 | neocloud |
| Lambda | $2.99 | neocloud |
| DataCrunch | $3.02 | neocloud |
| Crusoe | $3.10 | neocloud |
| Together | $3.15 | neocloud |
| DigitalOcean | $3.44 | neocloud |
| Vast.ai | $3.91 | Verified-listing rate, refreshed daily |
| CoreWeave | $4.15 | neocloud |
| Modal | $4.54 | Serverless, per-second billing |
| RunPod | $4.59 | Secure Cloud |
| OCI | $5.20 | BM.GPU.H200.8 ÷ 8 |
| AWS | $6.20 | p5e.48xlarge ÷ 8 |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Hopper (2024) |
|---|---|
| Memory | 141 GB HBM3e, 4.8 TB/s |
| Interconnect | NVLink 4 · 900 GB/s |
| Form factor | SXM |
| Partitioning | MIG, up to 7 isolated instances |
What fits on one H200
Weights + ~2 GB runtime overhead against 127 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Llama 3.3 70B | 70.6B | 8-bit |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | FP16 |
| Qwen2.5 32B | 32.8B | FP16 |
| Qwen2.5 Coder 32B | 32.8B | FP16 |
| Qwen2.5 72B | 72.7B | 8-bit |
| QwQ 32B (reasoning) | 32.8B | FP16 |
| Mistral 7B | 7.25B | FP16 |
| Mixtral 8x7B | 46.7B | FP16 |
| Mixtral 8x22B | 140.6B | 4-bit |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | FP16 |
| Phi-4 14B | 14.7B | FP16 |
| gpt-oss-20b | 20.9B | FP16 |
| gpt-oss-120b | 116.8B | 4-bit |
Best for
- 70B+ model inference
- Long-context serving (big KV caches)
- Training when memory per GPU is the constraint
Compare
H100 SXM vs H200
The workhorse of the AI boom. Still the most widely available serious training and inference GPU.
H200 vs B200
NVIDIA's Blackwell flagship: 192 GB of HBM3e and roughly double Hopper's throughput per chip.
H200 vs MI300X
AMD's answer to Hopper: 192 GB on one GPU, a 70B model in FP16 fits with room to spare.
H200 vs Gaudi 3
Intel's price-performance play, with Ethernet-native scale-out instead of proprietary interconnect.
H200 vs GH200
A Hopper GPU welded to a Grace CPU: 576 GB of unified fast memory for models that spill past VRAM.