RTX 4090 vs RTX 5090

RTX 5090 has 1.3× the memory; RTX 5090 has 1.8× the bandwidth; RTX 4090 rents for 1.6× less.

The NVIDIA RTX 4090 (Ada Lovelace, 2022) and the NVIDIA RTX 5090 (Blackwell, 2025) are 3 years and a silicon generation apart, which makes the price gap the heart of the story. On capacity, the RTX 5090 carries 1.3x the memory (32 vs 24 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the RTX 5090 leads at 1.79 vs 1.01 TB/s.

Price is where it settles: the RTX 4090 rents from $0.34/hr against $0.54/hr for the RTX 5090, a 1.6x gap that the performance numbers above do justify for the right workload. Per gigabyte of VRAM, the RTX 4090 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. Bear in mind the RTX 4090 is a consumer card rented mostly through marketplaces, with the trust and reliability trade-offs that implies, while the RTX 5090 is datacenter silicon with ECC HBM and standard hosting. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

NVIDIA · Ada Lovelace · 2022

NVIDIA RTX 4090

The people's inference GPU. On marketplaces, the best $/FLOP in the catalog.

From $0.34/hr at RunPod

NVIDIA · Blackwell · 2025

NVIDIA RTX 5090

The fastest thing you can put in a desktop: 32 GB of GDDR7 at 1.8 TB/s.

From $0.54/hr at Vast.ai

Head to head

■ RTX 4090   ■ RTX 5090, bars share one scale across the whole catalog.

Memory 24 GB 32 GB
Memory bandwidth 1.01 TB/s 1.79 TB/s
FP16 dense 165 TFLOPS 210 TFLOPS
FP8 dense 330 TFLOPS 420 TFLOPS
Power (TDP) 450 W 575 W
RTX 4090RTX 5090
Memory24 GB GDDR6X32 GB GDDR7
Bandwidth1.01 TB/s1.79 TB/s
FP16 dense165 TF210 TF
FP8 dense330 TF420 TF
TDP450 W575 W
InterconnectPCIe Gen4PCIe Gen5
Cheapest rental $0.34/hr $0.54/hr
$/hr per GB VRAM $1.4¢ $1.7¢

What one can run that the other can't

Only on RTX 4090

Nothing, RTX 5090 runs everything RTX 4090 does.

Only on RTX 5090

  • Qwen2.5 32B (4-bit)
  • Qwen2.5 Coder 32B (4-bit)
  • QwQ 32B (reasoning) (4-bit)

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.