NVIDIA · Ampere · 2020
NVIDIA A100 40GB
Same silicon as the 80GB at roughly half the price, the best cost-per-token machine in the datacenter tier when the model fits.
Identical GA100 compute, NVLink, and MIG support as the A100 80GB; the trade is 40 GB of HBM2 at 1.56 TB/s. Per rental dollar it delivers more bandwidth than the 80GB, so for the high-volume 7B-14B FP16 tier, 32B-class 4-bit serving, MIG-sliced endpoint fleets, and batch pipelines it is the sharper buy. Two 40GBs cost about one 80GB and out-serve it on a tensor-parallel 4-bit 70B.
Deep dive: A100 40GB vs 80GB, when half the price wins →
Where to rent a A100 40GB
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: A100 40GB pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| Vast.ai | $0.48 | Verified-listing rate, refreshed daily |
| Thunder | $0.66 | Virtualized instances |
| Jarvislabs | $0.89 | neocloud |
| RunPod | $0.99 | neocloud |
| GCP | $2.93 | a2-highgpu ÷ 8 |
| OCI | $3.05 | hyperscaler |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Ampere (2020) |
|---|---|
| Memory | 40 GB HBM2, 1.56 TB/s |
| Interconnect | NVLink 3 · 600 GB/s |
| Form factor | SXM / PCIe |
| Partitioning | MIG, up to 7 isolated instances |
What fits on one A100 40GB
Weights + ~2 GB runtime overhead against 36 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | FP16 |
| Qwen2.5 32B | 32.8B | 4-bit |
| Qwen2.5 Coder 32B | 32.8B | 4-bit |
| QwQ 32B (reasoning) | 32.8B | 4-bit |
| Mistral 7B | 7.25B | FP16 |
| Mixtral 8x7B | 46.7B | 4-bit |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | 8-bit |
| Phi-4 14B | 14.7B | FP16 |
| gpt-oss-20b | 20.9B | 8-bit |
Best for
- 7B-14B FP16 production serving
- 32B-class 4-bit inference and QLoRA
- MIG fleets: up to 7 isolated endpoints per card
- Batch pipelines: embeddings, evals, extraction