RTX 4090 vs RTX 5090
RTX 5090 has 1.3× the memory; RTX 5090 has 1.8× the bandwidth; RTX 4090 rents for 1.6× less.
The NVIDIA RTX 4090 (Ada Lovelace, 2022) and the NVIDIA RTX 5090 (Blackwell, 2025) are 3 years and a silicon generation apart, which makes the price gap the heart of the story. On capacity, the RTX 5090 carries 1.3x the memory (32 vs 24 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the RTX 5090 leads at 1.79 vs 1.01 TB/s.
Price is where it settles: the RTX 4090 rents from $0.34/hr against $0.54/hr for the RTX 5090, a 1.6x gap that the performance numbers above do justify for the right workload. Per gigabyte of VRAM, the RTX 4090 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. Bear in mind the RTX 4090 is a consumer card rented mostly through marketplaces, with the trust and reliability trade-offs that implies, while the RTX 5090 is datacenter silicon with ECC HBM and standard hosting. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.
NVIDIA · Ada Lovelace · 2022
NVIDIA RTX 4090
The people's inference GPU. On marketplaces, the best $/FLOP in the catalog.
From $0.34/hr at RunPod
NVIDIA · Blackwell · 2025
NVIDIA RTX 5090
The fastest thing you can put in a desktop: 32 GB of GDDR7 at 1.8 TB/s.
From $0.54/hr at Vast.ai
Head to head
■ RTX 4090 ■ RTX 5090, bars share one scale across the whole catalog.
| RTX 4090 | RTX 5090 | |
|---|---|---|
| Memory | 24 GB GDDR6X | 32 GB GDDR7 |
| Bandwidth | 1.01 TB/s | 1.79 TB/s |
| FP16 dense | 165 TF | 210 TF |
| FP8 dense | 330 TF | 420 TF |
| TDP | 450 W | 575 W |
| Interconnect | PCIe Gen4 | PCIe Gen5 |
| Cheapest rental | $0.34/hr | $0.54/hr |
| $/hr per GB VRAM | $1.4¢ | $1.7¢ |
What one can run that the other can't
Only on RTX 4090
Nothing, RTX 5090 runs everything RTX 4090 does.
Only on RTX 5090
- Qwen2.5 32B (4-bit)
- Qwen2.5 Coder 32B (4-bit)
- QwQ 32B (reasoning) (4-bit)
Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.