RTX 3090 vs RTX 4090

RTX 3090 rents for 1.9× less.

The NVIDIA RTX 3090 (Ampere, 2020) and the NVIDIA RTX 4090 (Ada Lovelace, 2022) are 2 years and a silicon generation apart, which makes the price gap the heart of the story. Memory is effectively a wash at 24 vs 24 GB, so capacity does not decide this one. Bandwidth is nearly even (0.936 vs 1.01 TB/s), so serving speed per GPU will be similar. Raw compute favors the RTX 4090 by 2.3x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding. The RTX 4090 also speaks native FP8 while the RTX 3090 tops out at BF16/INT8, a generational gap explained in CUDA cores vs Tensor Cores.

Price is where it settles: the RTX 3090 rents from $0.18/hr against $0.34/hr for the RTX 4090, a 1.9x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the RTX 3090 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. Bear in mind the RTX 3090 is a consumer card rented mostly through marketplaces, with the trust and reliability trade-offs that implies, while the RTX 4090 is datacenter silicon with ECC HBM and standard hosting. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

NVIDIA · Ampere · 2020

NVIDIA RTX 3090

The used-market classic: 24 GB for a few hundred dollars, or pennies per hour rented.

From $0.18/hr at Vast.ai

NVIDIA · Ada Lovelace · 2022

NVIDIA RTX 4090

The people's inference GPU. On marketplaces, the best $/FLOP in the catalog.

From $0.34/hr at RunPod

Head to head

■ RTX 3090   ■ RTX 4090, bars share one scale across the whole catalog.

Memory 24 GB 24 GB
Memory bandwidth 0.936 TB/s 1.01 TB/s
FP16 dense 71 TFLOPS 165 TFLOPS
FP8 dense - 330 TFLOPS
Power (TDP) 350 W 450 W
RTX 3090RTX 4090
Memory24 GB GDDR6X24 GB GDDR6X
Bandwidth0.936 TB/s1.01 TB/s
FP16 dense71 TF165 TF
FP8 dense-330 TF
TDP350 W450 W
InterconnectPCIe Gen4 (NVLink bridge)PCIe Gen4
Cheapest rental $0.18/hr $0.34/hr
$/hr per GB VRAM $0.8¢ $1.4¢

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.