H200 vs B200
B200 has 1.4× the memory; B200 has 1.7× the bandwidth; H200 rents for 1.6× less.
The NVIDIA H200 (Hopper, 2024) and the NVIDIA B200 (Blackwell, 2024) are close contemporaries from adjacent generations, so the details decide. On capacity, the B200 carries 1.4x the memory (192 vs 141 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the B200 leads at 8 vs 4.8 TB/s. Raw compute favors the B200 by 2.3x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding.
Price is where it settles: the H200 rents from $2.30/hr against $3.75/hr for the B200, a 1.6x gap that the performance numbers above do justify for the right workload. Per gigabyte of VRAM, the H200 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.
NVIDIA · Hopper · 2024
NVIDIA H200
An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.
From $2.30/hr at Nebius
NVIDIA · Blackwell · 2024
NVIDIA B200
NVIDIA's Blackwell flagship: 192 GB of HBM3e and roughly double Hopper's throughput per chip.
From $3.75/hr at Packet.ai
Head to head
■ H200 ■ B200, bars share one scale across the whole catalog.
| H200 | B200 | |
|---|---|---|
| Memory | 141 GB HBM3e | 192 GB HBM3e |
| Bandwidth | 4.8 TB/s | 8 TB/s |
| FP16 dense | 990 TF | 2,250 TF |
| FP8 dense | 1,979 TF | 4,500 TF |
| TDP | 700 W | 1000 W |
| Interconnect | NVLink 4 · 900 GB/s | NVLink 5 · 1.8 TB/s |
| Cheapest rental | $2.30/hr | $3.75/hr |
| $/hr per GB VRAM | $1.6¢ | $2.0¢ |
Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.