NVIDIA · Blackwell · 2024
NVIDIA B200
NVIDIA's Blackwell flagship: 192 GB of HBM3e and roughly double Hopper's throughput per chip.
Dual-die Blackwell package. FP4 support with second-generation Transformer Engine makes it the current default for frontier-scale training and high-throughput inference. Usually sold as 8-GPU HGX B200 systems.
Where to rent a B200
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: B200 pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| Packet.ai | $3.75 | Dynamic (shared multi-tenant) tier; dedicated $6.99 |
| Nebius | $4.80 | neocloud |
| Lambda | $4.99 | neocloud |
| Crusoe | $5.20 | neocloud |
| Together | $5.50 | neocloud |
| Modal | $6.25 | Serverless, per-second billing |
| CoreWeave | $6.48 | HGX B200 on-demand; reserved much lower |
| RunPod | $6.79 | neocloud |
| GCP | $8.20 | A4 ÷ 8 |
| AWS | $8.65 | p6-b200.48xlarge ÷ 8 |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Blackwell (2024) |
|---|---|
| Memory | 192 GB HBM3e, 8 TB/s |
| Interconnect | NVLink 5 · 1.8 TB/s |
| Form factor | SXM |
| Partitioning | MIG, up to 7 isolated instances |
What fits on one B200
Weights + ~2 GB runtime overhead against 173 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Llama 3.3 70B | 70.6B | FP16 |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | FP16 |
| Qwen2.5 32B | 32.8B | FP16 |
| Qwen2.5 Coder 32B | 32.8B | FP16 |
| Qwen2.5 72B | 72.7B | FP16 |
| QwQ 32B (reasoning) | 32.8B | FP16 |
| Mistral 7B | 7.25B | FP16 |
| Mixtral 8x7B | 46.7B | FP16 |
| Mixtral 8x22B | 140.6B | 8-bit |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | FP16 |
| Phi-4 14B | 14.7B | FP16 |
| gpt-oss-20b | 20.9B | FP16 |
| gpt-oss-120b | 116.8B | 8-bit |
Best for
- Frontier model training
- High-throughput FP8/FP4 inference
- Models up to ~180B params on a single GPU (4-bit)
Compare
H100 SXM vs B200
The workhorse of the AI boom. Still the most widely available serious training and inference GPU.
H200 vs B200
An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.
B200 vs MI325X
256 GB of HBM3e, the largest memory pool on any single accelerator you can rent.