The A100 in 2026: the best value in AI inference
12 min read · updated 2026-09
The A100 launched in May 2020, trained the models that started this era, and was declared obsolete by three successive NVIDIA generations. It is also, in 2026, arguably the best price-performance GPU you can point at an inference workload. Both things are true, and the second one is just arithmetic.
The only number that matters: bandwidth per dollar
LLM inference is memory-bound: generating each token requires streaming the model's active weights through the GPU, so tokens/second tracks memory bandwidth, not FLOPS (the full mechanics). That reframes the generational comparison entirely:
| GPU | Bandwidth | Cheapest rent* | TB/s per $/hr | FP16 TF per $/hr |
|---|---|---|---|---|
| A100 40GB | 1.56 TB/s | $0.55 | 2.84 | 567 |
| A100 80GB | 2.0 TB/s | $0.85 | 2.35 | 367 |
| H100 SXM | 3.35 TB/s | $1.65 | 2.03 | 600 |
| H200 | 4.8 TB/s | $2.10 | 2.29 | 471 |
| B200 | 8.0 TB/s | $4.80 | 1.67 | 469 |
*Cheapest tracked on-demand price, 2026-09 (methodology). Yes, the H100's FLOPS-per-dollar is better, which is why the H100 wins for training and prefill-heavy work, and the A100 wins for decode-heavy serving.
Per dollar of rent, the six-year-old card moves more bytes than anything NVIDIA has shipped since, except the H200 at its best price, and the A100's own best price keeps falling as hyperscalers and AI labs rotate fleets to Blackwell. Depreciation did what engineering couldn't: it made Ampere the value king. Note who tops the table: the 40GB variant. Both A100s share the same GA100 silicon and 312 TF of compute; the 40GB gives up 22% of the bandwidth and half the memory for roughly half the price, which makes it the better buy whenever the workload fits, a case worked in full in A100 40GB vs 80GB.
What an A100 actually serves in 2026
- Llama 3.3 / Qwen 2.5 70B-class at 4-bit on a single 80 GB card, ~45 GB of weights, real KV headroom, ~40 tok/s single-stream and far more batched. This is the workhorse configuration of the open-model economy.
- 7B-14B models in FP16 at scale, with 80 GB you serve them with huge batches, or slice the card into seven MIG instances and run seven isolated small-model endpoints on one device.
- Fine-tuning with LoRA/QLoRA, a single A100 fine-tunes a 4-bit 70B; a pair does it comfortably. Full-parameter training of small models still fits its 312 TF of BF16.
- Embeddings, rerankers, Whisper, diffusion, massively over-served by an A100, which is the point: at $0.85/hr you can afford to over-serve.
What it can't do: FP8 (Ampere predates Transformer Engine), FP4, or frontier-scale training economics. If your workload is training beyond LoRA scale or ultra-low-latency 70B+ serving, spend up, the A100 vs H100 decision guide draws the exact line.
The supply story: your gain is a hyperscaler's rotation
Why is the best value GPU six years old? Because the fleets that bought Ampere in 2020-22, hyperscalers, AI labs, autonomous-vehicle programs, are rotating to Hopper and Blackwell, and hundreds of thousands of working A100s are flowing into the secondary market and value-tier clouds. An accelerator built for a 5+ year datacenter duty cycle doesn't stop working because a faster one exists; HBM doesn't wear out on a marketing schedule. The result is a two-tier market:
- Rental: $0.85-1.40/hr at marketplaces and neoclouds, while the same card lists at $3.67-4.10/hr at hyperscalers that haven't repriced their Ampere fleets.
- Ownership: ~$6,800 for a used 80 GB card (~$3,400 for 40 GB), against a launch price north of $15,000. At 60% utilization and industrial power, that's an effective ~$0.35/GPU-hr, run your own numbers in the ROI calculator.
This is the classic trailing-edge compute story: the same dynamic that keeps 28nm fabs profitable and mainframes running keeps Ampere earning. Inference demand is growing faster than any single generation of supply, and most inference tokens don't need the newest silicon, they need the cheapest bandwidth that clears the latency bar. For a deeper look at where retired fleets go, see the second life of datacenter GPUs.
When to skip the A100
Honesty makes this case stronger, so: choose newer silicon when you need FP8/FP4 throughput (2× effective bandwidth from quantized weights on Hopper/Blackwell), sub-second first-token latency on 70B+ at high concurrency, 128K+ contexts (KV cache eats the A100's headroom, the H200 exists for this), or dense training beyond fine-tuning. And at the small end, a used RTX 4090 undercuts even the A100 for ≤13B hobby workloads.
The bottom line
Buy, or rent, the A100 for what it is in 2026: not a compromise, but the market's best-priced memory bandwidth wearing last generation's badge. Check what it can run against your model in the VRAM calculator, price the token economics in the inference cost estimator, and if you're weighing buying pulled fleet units outright, the ROI calculator will tell you where break-even sits.