L40S vs RTX 4090

L40S has 2.0× the memory; RTX 4090 has 1.2× the bandwidth; RTX 4090 rents for 2.4× less.

The NVIDIA L40S (Ada Lovelace, 2023) and the NVIDIA RTX 4090 (Ada Lovelace, 2022) are an odd couple on paper, the L40S is datacenter silicon and the RTX 4090 is a consumer flagship, and the price gap is exactly why people cross-shop them. On capacity, the L40S carries 2.0x the memory (48 vs 24 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the RTX 4090 leads at 1.01 vs 0.864 TB/s. Raw compute favors the L40S by 2.2x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding.

Price is where it settles: the RTX 4090 rents from $0.34/hr against $0.80/hr for the L40S, a 2.4x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the RTX 4090 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. Bear in mind the RTX 4090 is a consumer card rented mostly through marketplaces, with the trust and reliability trade-offs that implies, while the L40S is datacenter silicon with ECC HBM and standard hosting. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

NVIDIA · Ada Lovelace · 2023

NVIDIA L40S

Ada silicon with 48 GB: strong FP8 compute, modest bandwidth. A favorite for image gen and mid-size LLMs.

From $0.80/hr at Vast.ai

NVIDIA · Ada Lovelace · 2022

NVIDIA RTX 4090

The people's inference GPU. On marketplaces, the best $/FLOP in the catalog.

From $0.34/hr at RunPod

Head to head

■ L40S   ■ RTX 4090, bars share one scale across the whole catalog.

Memory 48 GB 24 GB
Memory bandwidth 0.864 TB/s 1.01 TB/s
FP16 dense 362 TFLOPS 165 TFLOPS
FP8 dense 733 TFLOPS 330 TFLOPS
Power (TDP) 350 W 450 W
L40SRTX 4090
Memory48 GB GDDR624 GB GDDR6X
Bandwidth0.864 TB/s1.01 TB/s
FP16 dense362 TF165 TF
FP8 dense733 TF330 TF
TDP350 W450 W
InterconnectPCIe Gen4 · 64 GB/sPCIe Gen4
Cheapest rental $0.80/hr $0.34/hr
$/hr per GB VRAM $1.7¢ $1.4¢

What one can run that the other can't

Only on L40S

  • Qwen2.5 32B (8-bit)
  • Qwen2.5 Coder 32B (8-bit)
  • QwQ 32B (reasoning) (8-bit)
  • Mixtral 8x7B (4-bit)

Only on RTX 4090

Nothing, L40S runs everything RTX 4090 does.

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.