NVIDIA · Ampere · 2021

NVIDIA A10

The quiet default for 7B-class production inference on AWS and Oracle.

24 GB
GDDR6
0.6 TB/s
Mem bandwidth
125 TF
FP16 dense
150 W
TDP
$0.75 /hr
From · Lambda

AWS g5 instances made the A10G ubiquitous for small-model serving. 24 GB comfortably fits a 13B at 8-bit or 7B at FP16 with real batch sizes.

Where to rent a A10

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: A10 pricing.

Provider $/GPU-hr vs cheapest Notes
Lambda $0.75 cheapest neocloud
AWS $1.01 1.3× g5.xlarge (A10G)
Modal $1.10 1.5× Serverless A10G
OCI $2.00 2.7× VM.GPU.A10.1

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 24 GB
Memory bandwidth 0.6 TB/s
FP16 dense 125 TFLOPS
FP8 dense -
Power (TDP) 150 W
ArchitectureAmpere (2021)
Memory24 GB GDDR6, 0.6 TB/s
InterconnectPCIe Gen4
Form factorPCIe
PartitioningNo MIG (time-slicing / vGPU only)

What fits on one A10

Weights + ~2 GB runtime overhead against 22 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B 8-bit
Mistral 7B 7.25B FP16
Gemma 2 9B 9.24B 8-bit
Gemma 2 27B 27.2B 4-bit
Phi-4 14B 14.7B 8-bit
gpt-oss-20b 20.9B 4-bit

Best for

Compare

A10 vs L4

72 watts. The efficiency play for video, small-model inference, and high-density serving.

A10 vs T4

The 2018 inference stalwart. Slow by modern standards, but everywhere and nearly free.