NVIDIA · Ampere · 2020

NVIDIA A100 80GB

The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.

80 GB
HBM2e
2 TB/s
Mem bandwidth
312 TF
FP16 dense
400 W
TDP
$0.87 /hr
From · Vast.ai

No FP8 support (Ampere predates Transformer Engine), so it runs FP16/BF16 or INT8. At 2026 rental prices it is often the best $/token for 7B-70B class models where H100 speed isn't needed.

Deep dive: The A100 in 2026, the best value in AI inference →

Where to rent a A100 80GB

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: A100 80GB pricing.

Provider $/GPU-hr vs cheapest Notes
Vast.ai $0.87 cheapest Verified-listing rate, refreshed daily
Thunder $1.09 1.3× Virtualized instances
Nebius $1.15 1.3× neocloud
DataCrunch $1.15 1.3× neocloud
TensorDock $1.20 1.4× marketplace
Lambda $1.29 1.5× neocloud
Hyperstack $1.35 1.6× neocloud
Massed $1.35 1.6× neocloud
Cudo $1.35 1.6× marketplace
RunPod $1.39 1.6× Secure Cloud; Community from ~$1.09
Packet.ai $1.43 1.6× No-contract on-demand; monthly from $940
DigitalOcean $1.89 2.2× neocloud
CoreWeave $2.06 2.4× neocloud
Modal $2.50 2.9× Serverless, per-second billing
Azure $3.67 4.2× ND A100 v4 ÷ 8
GCP $3.93 4.5× a2-ultragpu ÷ 8
AWS $4.10 4.7× p4de.24xlarge ÷ 8

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 80 GB
Memory bandwidth 2 TB/s
FP16 dense 312 TFLOPS
FP8 dense -
Power (TDP) 400 W
ArchitectureAmpere (2020)
Memory80 GB HBM2e, 2 TB/s
InterconnectNVLink 3 · 600 GB/s
Form factorSXM / PCIe
PartitioningMIG, up to 7 isolated instances

What fits on one A100 80GB

Weights + ~2 GB runtime overhead against 72 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Llama 3.3 70B 70.6B 4-bit
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B FP16
Qwen2.5 32B 32.8B 8-bit
Qwen2.5 Coder 32B 32.8B 8-bit
Qwen2.5 72B 72.7B 4-bit
QwQ 32B (reasoning) 32.8B 8-bit
Mistral 7B 7.25B FP16
Mixtral 8x7B 46.7B 8-bit
Gemma 2 9B 9.24B FP16
Gemma 2 27B 27.2B FP16
Phi-4 14B 14.7B FP16
gpt-oss-20b 20.9B FP16

Best for

Compare

H100 SXM vs A100 80GB

The workhorse of the AI boom. Still the most widely available serious training and inference GPU.

A100 80GB vs A100 40GB

Same silicon as the 80GB at roughly half the price, the best cost-per-token machine in the datacenter tier when the model fits.

A100 80GB vs L40S

Ada silicon with 48 GB: strong FP8 compute, modest bandwidth. A favorite for image gen and mid-size LLMs.

A100 80GB vs RTX 4090

The people's inference GPU. On marketplaces, the best $/FLOP in the catalog.

A100 80GB vs RTX 5090

The fastest thing you can put in a desktop: 32 GB of GDDR7 at 1.8 TB/s.

Gaudi 3 vs A100 80GB

Intel's price-performance play, with Ethernet-native scale-out instead of proprietary interconnect.

A100 80GB vs V100

The first Tensor Core GPU. Retired from the frontier, still cheap and capable for small models.