NVIDIA · Ampere · 2020

NVIDIA A100 40GB

Same silicon as the 80GB at roughly half the price, the best cost-per-token machine in the datacenter tier when the model fits.

40 GB
HBM2
1.56 TB/s
Mem bandwidth
312 TF
FP16 dense
400 W
TDP
$0.48 /hr
From · Vast.ai

Identical GA100 compute, NVLink, and MIG support as the A100 80GB; the trade is 40 GB of HBM2 at 1.56 TB/s. Per rental dollar it delivers more bandwidth than the 80GB, so for the high-volume 7B-14B FP16 tier, 32B-class 4-bit serving, MIG-sliced endpoint fleets, and batch pipelines it is the sharper buy. Two 40GBs cost about one 80GB and out-serve it on a tensor-parallel 4-bit 70B.

Deep dive: A100 40GB vs 80GB, when half the price wins →

Where to rent a A100 40GB

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: A100 40GB pricing.

Provider $/GPU-hr vs cheapest Notes
Vast.ai $0.48 cheapest Verified-listing rate, refreshed daily
Thunder $0.66 1.4× Virtualized instances
Jarvislabs $0.89 1.9× neocloud
RunPod $0.99 2.1× neocloud
GCP $2.93 6.1× a2-highgpu ÷ 8
OCI $3.05 6.4× hyperscaler

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 40 GB
Memory bandwidth 1.56 TB/s
FP16 dense 312 TFLOPS
FP8 dense -
Power (TDP) 400 W
ArchitectureAmpere (2020)
Memory40 GB HBM2, 1.56 TB/s
InterconnectNVLink 3 · 600 GB/s
Form factorSXM / PCIe
PartitioningMIG, up to 7 isolated instances

What fits on one A100 40GB

Weights + ~2 GB runtime overhead against 36 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B FP16
Qwen2.5 32B 32.8B 4-bit
Qwen2.5 Coder 32B 32.8B 4-bit
QwQ 32B (reasoning) 32.8B 4-bit
Mistral 7B 7.25B FP16
Mixtral 8x7B 46.7B 4-bit
Gemma 2 9B 9.24B FP16
Gemma 2 27B 27.2B 8-bit
Phi-4 14B 14.7B FP16
gpt-oss-20b 20.9B 8-bit

Best for

Compare

A100 80GB vs A100 40GB

The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.

A100 40GB vs RTX 5090

The fastest thing you can put in a desktop: 32 GB of GDDR7 at 1.8 TB/s.