L4 vs T4

L4 has 1.5× the memory; T4 rents for 3.2× less.

The NVIDIA L4 (Ada Lovelace, 2023) and the NVIDIA T4 (Turing, 2018) are 5 years and a silicon generation apart, which makes the price gap the heart of the story. On capacity, the L4 carries 1.5x the memory (24 vs 16 GB), which sets what each can hold at all. Bandwidth is nearly even (0.3 vs 0.32 TB/s), so serving speed per GPU will be similar. Raw compute favors the L4 by 1.9x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding. The L4 also speaks native FP8 while the T4 tops out at FP16, a generational gap explained in CUDA cores vs Tensor Cores.

Price is where it settles: the T4 rents from $0.10/hr against $0.32/hr for the L4, a 3.2x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the T4 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

NVIDIA · Ada Lovelace · 2023

NVIDIA L4

72 watts. The efficiency play for video, small-model inference, and high-density serving.

From $0.32/hr at Vast.ai

NVIDIA · Turing · 2018

NVIDIA T4

The 2018 inference stalwart. Slow by modern standards, but everywhere and nearly free.

From $0.10/hr at Vast.ai

Head to head

■ L4   ■ T4, bars share one scale across the whole catalog.

Memory 24 GB 16 GB
Memory bandwidth 0.3 TB/s 0.32 TB/s
FP16 dense 121 TFLOPS 65 TFLOPS
FP8 dense 242 TFLOPS -
Power (TDP) 72 W 70 W
L4T4
Memory24 GB GDDR616 GB GDDR6
Bandwidth0.3 TB/s0.32 TB/s
FP16 dense121 TF65 TF
FP8 dense242 TF-
TDP72 W70 W
InterconnectPCIe Gen4PCIe Gen3
Cheapest rental $0.32/hr $0.10/hr
$/hr per GB VRAM $1.3¢ $0.6¢

What one can run that the other can't

Only on L4

  • Gemma 2 27B (4-bit)
  • gpt-oss-20b (4-bit)

Only on T4

Nothing, L4 runs everything T4 does.

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.