A10 vs L4
A10 has 2.0× the bandwidth; L4 rents for 2.3× less.
The NVIDIA A10 (Ampere, 2021) and the NVIDIA L4 (Ada Lovelace, 2023) are 2 years and a silicon generation apart, which makes the price gap the heart of the story. Memory is effectively a wash at 24 vs 24 GB, so capacity does not decide this one. On memory bandwidth, the number that governs LLM serving speed, the A10 leads at 0.6 vs 0.3 TB/s. The L4 also speaks native FP8 while the A10 tops out at BF16/INT8, a generational gap explained in CUDA cores vs Tensor Cores.
Price is where it settles: the L4 rents from $0.32/hr against $0.75/hr for the A10, a 2.3x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the L4 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.
NVIDIA · Ampere · 2021
NVIDIA A10
The quiet default for 7B-class production inference on AWS and Oracle.
From $0.75/hr at Lambda
Head to head
■ A10 ■ L4, bars share one scale across the whole catalog.
| A10 | L4 | |
|---|---|---|
| Memory | 24 GB GDDR6 | 24 GB GDDR6 |
| Bandwidth | 0.6 TB/s | 0.3 TB/s |
| FP16 dense | 125 TF | 121 TF |
| FP8 dense | - | 242 TF |
| TDP | 150 W | 72 W |
| Interconnect | PCIe Gen4 | PCIe Gen4 |
| Cheapest rental | $0.75/hr | $0.32/hr |
| $/hr per GB VRAM | $3.1¢ | $1.3¢ |
Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.