A100 80GB vs V100

A100 80GB has 2.5× the memory; A100 80GB has 2.2× the bandwidth; V100 rents for 4.1× less.

The NVIDIA A100 80GB (Ampere, 2020) and the NVIDIA V100 (Volta, 2017) are 3 years and a silicon generation apart, which makes the price gap the heart of the story. On capacity, the A100 80GB carries 2.5x the memory (80 vs 32 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the A100 80GB leads at 2 vs 0.9 TB/s. Raw compute favors the A100 80GB by 2.5x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding.

Price is where it settles: the V100 rents from $0.21/hr against $0.87/hr for the A100 80GB, a 4.1x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the V100 is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For why Ampere remains the value benchmark this comparison is priced against, see The A100 in 2026. Everything below is the underlying data.

NVIDIA · Ampere · 2020

NVIDIA A100 80GB

The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.

From $0.87/hr at Vast.ai

NVIDIA · Volta · 2017

NVIDIA V100

The first Tensor Core GPU. Retired from the frontier, still cheap and capable for small models.

From $0.21/hr at Vast.ai

Head to head

■ A100 80GB   ■ V100, bars share one scale across the whole catalog.

Memory 80 GB 32 GB
Memory bandwidth 2 TB/s 0.9 TB/s
FP16 dense 312 TFLOPS 125 TFLOPS
FP8 dense - -
Power (TDP) 400 W 300 W
A100 80GBV100
Memory80 GB HBM2e32 GB HBM2
Bandwidth2 TB/s0.9 TB/s
FP16 dense312 TF125 TF
FP8 dense--
TDP400 W300 W
InterconnectNVLink 3 · 600 GB/sNVLink 2 · 300 GB/s
Cheapest rental $0.87/hr $0.21/hr
$/hr per GB VRAM $1.1¢ $0.7¢

What one can run that the other can't

Only on A100 80GB

  • Llama 3.3 70B (4-bit)
  • Qwen2.5 72B (4-bit)
  • Mixtral 8x7B (8-bit)

Only on V100

Nothing, A100 80GB runs everything V100 does.

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.