A100 80GB vs A100 40GB

A100 80GB has 2.0× the memory; A100 80GB has 1.3× the bandwidth; A100 40GB rents for 1.8× less.

The NVIDIA A100 80GB (Ampere, 2020) and the NVIDIA A100 40GB (Ampere, 2020) share the same Ampere silicon generation, so this comes down to configuration and price rather than architecture. On capacity, the A100 80GB carries 2.0x the memory (80 vs 40 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the A100 80GB leads at 2 vs 1.56 TB/s.

Price is where it settles: the A100 40GB rents from $0.48/hr against $0.87/hr for the A100 80GB, a 1.8x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the A100 80GB is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. The full decision guide for this exact pair is in A100 40GB vs 80GB: when half the price wins. Everything below is the underlying data.

NVIDIA · Ampere · 2020

NVIDIA A100 80GB

The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.

From $0.87/hr at Vast.ai

NVIDIA · Ampere · 2020

NVIDIA A100 40GB

Same silicon as the 80GB at roughly half the price, the best cost-per-token machine in the datacenter tier when the model fits.

From $0.48/hr at Vast.ai

Head to head

■ A100 80GB   ■ A100 40GB, bars share one scale across the whole catalog.

Memory 80 GB 40 GB
Memory bandwidth 2 TB/s 1.56 TB/s
FP16 dense 312 TFLOPS 312 TFLOPS
FP8 dense - -
Power (TDP) 400 W 400 W
A100 80GBA100 40GB
Memory80 GB HBM2e40 GB HBM2
Bandwidth2 TB/s1.56 TB/s
FP16 dense312 TF312 TF
FP8 dense--
TDP400 W400 W
InterconnectNVLink 3 · 600 GB/sNVLink 3 · 600 GB/s
Cheapest rental $0.87/hr $0.48/hr
$/hr per GB VRAM $1.1¢ $1.2¢

What one can run that the other can't

Only on A100 80GB

  • Llama 3.3 70B (4-bit)
  • Qwen2.5 72B (4-bit)

Only on A100 40GB

Nothing, A100 80GB runs everything A100 40GB does.

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.