A100 40GB vs 80GB: when half the price wins
9 min read · updated 2026-09
The A100 comes in two memory sizes built on the identical chip, and the market prices them a factor of two apart. That gap is not a verdict on the 40GB, it's an arbitrage, because for a large class of inference work the cards are interchangeable and one of them costs half as much.
Same engine, smaller tank
| A100 40GB | A100 80GB | |
|---|---|---|
| Silicon / compute | Identical (GA100 · 312 TF) | |
| MIG / NVLink | Identical (×7 · 600 GB/s) | |
| Memory | 40 GB HBM2 | 80 GB HBM2e |
| Bandwidth | 1.56 TB/s | 2.0 TB/s |
| Cheapest rent (2026-09) | $0.55/hr | $0.85/hr |
| Typical secure-cloud rent | $0.99/hr | $1.39/hr |
| Indicative unit price | ~$3,400 | ~$6,800 |
| Bandwidth per rental $ | 2.84 TB/s per $/hr | 2.35 TB/s per $/hr |
Read the last row again: per dollar, the 40GB moves more bytes, and bytes moved per dollar is what inference cost is made of. The only question that matters is whether your workload fits in the tank. When it does, paying for the 80GB is paying for empty memory.
Where the 40GB is simply the right card
- The 7B-14B production tier. Llama 3.1 8B, Qwen 2.5 7B/14B, Phi-4 in FP16 with real batch depth, this is the highest-volume tier of the open-model economy (routed stacks send 80% of traffic here; see cost per token), and 40 GB serves it with room to spare.
- 32B-class at 4-bit. Qwen 2.5 32B / Coder / QwQ in ~19 GB of weights leaves workable KV headroom at moderate context.
- MIG fleets. The 40GB slices into up to seven isolated 5 GB instances, seven small-model or embedding endpoints per card at the lowest cost per endpoint of any datacenter GPU (how MIG works).
- Embeddings, rerankers, Whisper, classic ML, SDXL-class image gen. None of these come near 40 GB; all of them enjoy HBM bandwidth that no consumer-card farm matches for reliability.
- Fine-tuning below ~34B. QLoRA on a 32B fits; LoRA on 7-14B is comfortable. Half-price training hours are half-price.
- Batch pipelines, extraction, evals, synthetic data, scanning: short-context, high-batch, KV-light work where the 40GB's per-dollar bandwidth advantage lands in full.
The two-for-one move
The subtler win: at ~half price, a pair of 40GBs costs what one 80GB does, and the pair brings 624 TF of compute, 3.1 TB/s of aggregate bandwidth, the same 80 GB of total memory, and NVLink between them. Tensor-parallel over two cards, a 4-bit 70B serves with higher throughput than on a single 80GB (aggregate bandwidth wins; TP overhead over NVLink costs ~10-15%). Two cards also mean two failure domains and two independently schedulable resources when the big model isn't running. The costs are real but bounded: a second PCIe/SXM slot, slightly more power, and multi-GPU serving config, table stakes for any team already running vLLM.
Where you should pay for the 80GB
Honesty first, as always: choose the 80GB for single-GPU 70B serving (4-bit weights plus real KV cache wants ~50+ GB), long-context work (KV cache scales with context, the math), large-batch serving of 30B-class models, and fine-tuning above ~34B on one card. If your roadmap says "we'll serve a 70B on each single card next quarter," buy the 80GB now.
The verdict
Price the decision, don't vibe it: put your model through the VRAM calculator, if it fits in 36 GB usable, the 40GB is the value play, and the cost estimator and ROI calculator will show you what the saved dollars compound into at fleet scale. The 80GB gets the headlines; the 40GB, at half the price for the same silicon, is quietly the best cost-per-token machine in the entire datacenter tier.
Need this at rack scale rather than card scale? Charg, from the team behind Virtualized, operates dedicated single-tenant A100 capacity: hyperscaler-grade fleets, validated and characterized in-house, U.S. datacenters, InfiniBand fabric, deployable now with no allocation queue. Reserve capacity →