L40S vs L4

L40S has 2.0× the memory; L40S has 2.9× the bandwidth; L4 rents for 2.5× less.

The NVIDIA L40S (Ada Lovelace, 2023) and the NVIDIA L4 (Ada Lovelace, 2023) share the same Ada Lovelace silicon generation, so this comes down to configuration and price rather than architecture. On capacity, the L40S carries 2.0x the memory (48 vs 24 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the L40S leads at 0.864 vs 0.3 TB/s. Raw compute favors the L40S by 3.0x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding.

Price is where it settles: the L4 rents from $0.32/hr against $0.80/hr for the L40S, a 2.5x gap that the performance numbers above do justify for the right workload. Per gigabyte of VRAM, the L4 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

NVIDIA · Ada Lovelace · 2023

NVIDIA L40S

Ada silicon with 48 GB: strong FP8 compute, modest bandwidth. A favorite for image gen and mid-size LLMs.

From $0.80/hr at Vast.ai

NVIDIA · Ada Lovelace · 2023

NVIDIA L4

72 watts. The efficiency play for video, small-model inference, and high-density serving.

From $0.32/hr at Vast.ai

Head to head

■ L40S   ■ L4, bars share one scale across the whole catalog.

Memory 48 GB 24 GB
Memory bandwidth 0.864 TB/s 0.3 TB/s
FP16 dense 362 TFLOPS 121 TFLOPS
FP8 dense 733 TFLOPS 242 TFLOPS
Power (TDP) 350 W 72 W
L40SL4
Memory48 GB GDDR624 GB GDDR6
Bandwidth0.864 TB/s0.3 TB/s
FP16 dense362 TF121 TF
FP8 dense733 TF242 TF
TDP350 W72 W
InterconnectPCIe Gen4 · 64 GB/sPCIe Gen4
Cheapest rental $0.80/hr $0.32/hr
$/hr per GB VRAM $1.7¢ $1.3¢

What one can run that the other can't

Only on L40S

  • Qwen2.5 32B (8-bit)
  • Qwen2.5 Coder 32B (8-bit)
  • QwQ 32B (reasoning) (8-bit)
  • Mixtral 8x7B (4-bit)

Only on L4

Nothing, L40S runs everything L4 does.

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.