NVIDIA · Blackwell · 2025
NVIDIA RTX 5090
The fastest thing you can put in a desktop: 32 GB of GDDR7 at 1.8 TB/s.
Blackwell consumer flagship. FP4 support and near-A100 memory bandwidth make it a serious local-inference machine; datacenter deployment is barred by NVIDIA's driver EULA, but GPU marketplaces rent them anyway.
Where to rent a RTX 5090
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: RTX 5090 pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| Vast.ai | $0.54 | Verified-listing rate, refreshed daily |
| RunPod | $0.69 | Community Cloud |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Blackwell (2025) |
|---|---|
| Memory | 32 GB GDDR7, 1.79 TB/s |
| Interconnect | PCIe Gen5 |
| Form factor | Consumer card |
| Partitioning | No MIG (time-slicing / vGPU only) |
What fits on one RTX 5090
Weights + ~2 GB runtime overhead against 29 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | 8-bit |
| Qwen2.5 32B | 32.8B | 4-bit |
| Qwen2.5 Coder 32B | 32.8B | 4-bit |
| QwQ 32B (reasoning) | 32.8B | 4-bit |
| Mistral 7B | 7.25B | FP16 |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | 4-bit |
| Phi-4 14B | 14.7B | 8-bit |
| gpt-oss-20b | 20.9B | 8-bit |
Best for
- Local LLMs up to ~32B (4-bit)
- Image/video generation
- Cheapest strong FLOPS on marketplaces
Compare
RTX 4090 vs RTX 5090
The people's inference GPU. On marketplaces, the best $/FLOP in the catalog.
RTX 5090 vs H100 SXM
The workhorse of the AI boom. Still the most widely available serious training and inference GPU.
A100 80GB vs RTX 5090
The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.
A100 40GB vs RTX 5090
Same silicon as the 80GB at roughly half the price, the best cost-per-token machine in the datacenter tier when the model fits.