NVIDIA · Blackwell · 2025

NVIDIA RTX PRO 6000 Blackwell

Blackwell with 96 GB in a PCIe card: workstation silicon that quietly became a serious inference rental.

96 GB
GDDR7 ECC
1.79 TB/s
Mem bandwidth
250 TF
FP16 dense
500 TF
FP8 dense
600 W
TDP
$0.50 /hr
From · RunPod

The full Blackwell GB202 die with 96 GB of ECC GDDR7 at 1.79 TB/s, FP8 and FP4 support, and none of the consumer-card licensing baggage. Clouds increasingly rent it as a mid-tier inference card: near-5090 bandwidth, four times the memory headroom of a 4090, and MIG partitioning.

Where to rent a RTX PRO 6000

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: RTX PRO 6000 pricing.

Provider $/GPU-hr vs cheapest Notes
RunPod $0.50 cheapest Secure Cloud
Packet.ai $0.66 1.3× Dynamic (shared) tier
Vast.ai $1.28 2.6× Verified-listing rate, refreshed daily

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 96 GB
Memory bandwidth 1.79 TB/s
FP16 dense 250 TFLOPS
FP8 dense 500 TFLOPS
Power (TDP) 600 W
ArchitectureBlackwell (2025)
Memory96 GB GDDR7 ECC, 1.79 TB/s
InterconnectPCIe Gen5
Form factorWorkstation / server PCIe
PartitioningMIG, up to 4 isolated instances

What fits on one RTX PRO 6000

Weights + ~2 GB runtime overhead against 86 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Llama 3.3 70B 70.6B 8-bit
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B FP16
Qwen2.5 32B 32.8B FP16
Qwen2.5 Coder 32B 32.8B FP16
Qwen2.5 72B 72.7B 8-bit
QwQ 32B (reasoning) 32.8B FP16
Mistral 7B 7.25B FP16
Mixtral 8x7B 46.7B 8-bit
Gemma 2 9B 9.24B FP16
Gemma 2 27B 27.2B FP16
Phi-4 14B 14.7B FP16
gpt-oss-20b 20.9B FP16
gpt-oss-120b 116.8B 4-bit

Best for