A10 vs L4

A10 has 2.0× the bandwidth; L4 rents for 2.3× less.

The NVIDIA A10 (Ampere, 2021) and the NVIDIA L4 (Ada Lovelace, 2023) are 2 years and a silicon generation apart, which makes the price gap the heart of the story. Memory is effectively a wash at 24 vs 24 GB, so capacity does not decide this one. On memory bandwidth, the number that governs LLM serving speed, the A10 leads at 0.6 vs 0.3 TB/s. The L4 also speaks native FP8 while the A10 tops out at BF16/INT8, a generational gap explained in CUDA cores vs Tensor Cores.

Price is where it settles: the L4 rents from $0.32/hr against $0.75/hr for the A10, a 2.3x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the L4 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

NVIDIA · Ampere · 2021

NVIDIA A10

The quiet default for 7B-class production inference on AWS and Oracle.

From $0.75/hr at Lambda

NVIDIA · Ada Lovelace · 2023

NVIDIA L4

72 watts. The efficiency play for video, small-model inference, and high-density serving.

From $0.32/hr at Vast.ai

Head to head

■ A10   ■ L4, bars share one scale across the whole catalog.

Memory 24 GB 24 GB
Memory bandwidth 0.6 TB/s 0.3 TB/s
FP16 dense 125 TFLOPS 121 TFLOPS
FP8 dense - 242 TFLOPS
Power (TDP) 150 W 72 W
A10L4
Memory24 GB GDDR624 GB GDDR6
Bandwidth0.6 TB/s0.3 TB/s
FP16 dense125 TF121 TF
FP8 dense-242 TF
TDP150 W72 W
InterconnectPCIe Gen4PCIe Gen4
Cheapest rental $0.75/hr $0.32/hr
$/hr per GB VRAM $3.1¢ $1.3¢

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.