A100 40GB vs RTX 5090

A100 40GB has 1.3× the memory.

The NVIDIA A100 40GB (Ampere, 2020) and the NVIDIA RTX 5090 (Blackwell, 2025) are an odd couple on paper, the A100 40GB is datacenter silicon and the RTX 5090 is a consumer flagship, and the price gap is exactly why people cross-shop them. On capacity, the A100 40GB carries 1.3x the memory (40 vs 32 GB), which sets what each can hold at all. Bandwidth is nearly even (1.56 vs 1.79 TB/s), so serving speed per GPU will be similar. Raw compute favors the A100 40GB by 1.5x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding. The RTX 5090 also speaks native FP8 while the A100 40GB tops out at BF16/INT8, a generational gap explained in CUDA cores vs Tensor Cores.

Price is where it settles: the A100 40GB rents from $0.48/hr against $0.54/hr for the RTX 5090, close enough that price should not decide. Per gigabyte of VRAM, the A100 40GB is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. Bear in mind the RTX 5090 is a consumer card rented mostly through marketplaces, with the trust and reliability trade-offs that implies, while the A100 40GB is datacenter silicon with ECC HBM and standard hosting. For why Ampere remains the value benchmark this comparison is priced against, see The A100 in 2026. Everything below is the underlying data.

NVIDIA · Ampere · 2020

NVIDIA A100 40GB

Same silicon as the 80GB at roughly half the price, the best cost-per-token machine in the datacenter tier when the model fits.

From $0.48/hr at Vast.ai

NVIDIA · Blackwell · 2025

NVIDIA RTX 5090

The fastest thing you can put in a desktop: 32 GB of GDDR7 at 1.8 TB/s.

From $0.54/hr at Vast.ai

Head to head

■ A100 40GB   ■ RTX 5090, bars share one scale across the whole catalog.

Memory 40 GB 32 GB
Memory bandwidth 1.56 TB/s 1.79 TB/s
FP16 dense 312 TFLOPS 210 TFLOPS
FP8 dense - 420 TFLOPS
Power (TDP) 400 W 575 W
A100 40GBRTX 5090
Memory40 GB HBM232 GB GDDR7
Bandwidth1.56 TB/s1.79 TB/s
FP16 dense312 TF210 TF
FP8 dense-420 TF
TDP400 W575 W
InterconnectNVLink 3 · 600 GB/sPCIe Gen5
Cheapest rental $0.48/hr $0.54/hr
$/hr per GB VRAM $1.2¢ $1.7¢

What one can run that the other can't

Only on A100 40GB

  • Mixtral 8x7B (4-bit)

Only on RTX 5090

Nothing, A100 40GB runs everything RTX 5090 does.

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.