NVIDIA · Ampere · 2020
NVIDIA A100 80GB
The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.
No FP8 support (Ampere predates Transformer Engine), so it runs FP16/BF16 or INT8. At 2026 rental prices it is often the best $/token for 7B-70B class models where H100 speed isn't needed.
Deep dive: The A100 in 2026, the best value in AI inference →
Where to rent a A100 80GB
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: A100 80GB pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| Vast.ai | $0.87 | Verified-listing rate, refreshed daily |
| Thunder | $1.09 | Virtualized instances |
| Nebius | $1.15 | neocloud |
| DataCrunch | $1.15 | neocloud |
| TensorDock | $1.20 | marketplace |
| Lambda | $1.29 | neocloud |
| Hyperstack | $1.35 | neocloud |
| Massed | $1.35 | neocloud |
| Cudo | $1.35 | marketplace |
| RunPod | $1.39 | Secure Cloud; Community from ~$1.09 |
| Packet.ai | $1.43 | No-contract on-demand; monthly from $940 |
| DigitalOcean | $1.89 | neocloud |
| CoreWeave | $2.06 | neocloud |
| Modal | $2.50 | Serverless, per-second billing |
| Azure | $3.67 | ND A100 v4 ÷ 8 |
| GCP | $3.93 | a2-ultragpu ÷ 8 |
| AWS | $4.10 | p4de.24xlarge ÷ 8 |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Ampere (2020) |
|---|---|
| Memory | 80 GB HBM2e, 2 TB/s |
| Interconnect | NVLink 3 · 600 GB/s |
| Form factor | SXM / PCIe |
| Partitioning | MIG, up to 7 isolated instances |
What fits on one A100 80GB
Weights + ~2 GB runtime overhead against 72 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Llama 3.3 70B | 70.6B | 4-bit |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | FP16 |
| Qwen2.5 32B | 32.8B | 8-bit |
| Qwen2.5 Coder 32B | 32.8B | 8-bit |
| Qwen2.5 72B | 72.7B | 4-bit |
| QwQ 32B (reasoning) | 32.8B | 8-bit |
| Mistral 7B | 7.25B | FP16 |
| Mixtral 8x7B | 46.7B | 8-bit |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | FP16 |
| Phi-4 14B | 14.7B | FP16 |
| gpt-oss-20b | 20.9B | FP16 |
Best for
- Budget fine-tuning (LoRA/QLoRA)
- 7B-34B model serving
- Batch inference where latency is flexible
Compare
H100 SXM vs A100 80GB
The workhorse of the AI boom. Still the most widely available serious training and inference GPU.
A100 80GB vs A100 40GB
Same silicon as the 80GB at roughly half the price, the best cost-per-token machine in the datacenter tier when the model fits.
A100 80GB vs L40S
Ada silicon with 48 GB: strong FP8 compute, modest bandwidth. A favorite for image gen and mid-size LLMs.
A100 80GB vs RTX 4090
The people's inference GPU. On marketplaces, the best $/FLOP in the catalog.
A100 80GB vs RTX 5090
The fastest thing you can put in a desktop: 32 GB of GDDR7 at 1.8 TB/s.
Gaudi 3 vs A100 80GB
Intel's price-performance play, with Ethernet-native scale-out instead of proprietary interconnect.
A100 80GB vs V100
The first Tensor Core GPU. Retired from the frontier, still cheap and capable for small models.