A100 80GB vs L40S

A100 80GB has 1.7× the memory; A100 80GB has 2.3× the bandwidth.

The NVIDIA A100 80GB (Ampere, 2020) and the NVIDIA L40S (Ada Lovelace, 2023) are 3 years and a silicon generation apart, which makes the price gap the heart of the story. On capacity, the A100 80GB carries 1.7x the memory (80 vs 48 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the A100 80GB leads at 2 vs 0.864 TB/s. The L40S also speaks native FP8 while the A100 80GB tops out at BF16/INT8, a generational gap explained in CUDA cores vs Tensor Cores.

Price is where it settles: the L40S rents from $0.80/hr against $0.87/hr for the A100 80GB, close enough that price should not decide. Per gigabyte of VRAM, the A100 80GB is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For why Ampere remains the value benchmark this comparison is priced against, see The A100 in 2026. Everything below is the underlying data.

NVIDIA · Ampere · 2020

NVIDIA A100 80GB

The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.

From $0.87/hr at Vast.ai

NVIDIA · Ada Lovelace · 2023

NVIDIA L40S

Ada silicon with 48 GB: strong FP8 compute, modest bandwidth. A favorite for image gen and mid-size LLMs.

From $0.80/hr at Vast.ai

Head to head

■ A100 80GB   ■ L40S, bars share one scale across the whole catalog.

Memory 80 GB 48 GB
Memory bandwidth 2 TB/s 0.864 TB/s
FP16 dense 312 TFLOPS 362 TFLOPS
FP8 dense - 733 TFLOPS
Power (TDP) 400 W 350 W
A100 80GBL40S
Memory80 GB HBM2e48 GB GDDR6
Bandwidth2 TB/s0.864 TB/s
FP16 dense312 TF362 TF
FP8 dense-733 TF
TDP400 W350 W
InterconnectNVLink 3 · 600 GB/sPCIe Gen4 · 64 GB/s
Cheapest rental $0.87/hr $0.80/hr
$/hr per GB VRAM $1.1¢ $1.7¢

What one can run that the other can't

Only on A100 80GB

  • Llama 3.3 70B (4-bit)
  • Qwen2.5 72B (4-bit)

Only on L40S

Nothing, A100 80GB runs everything L40S does.

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.