NVIDIA · Ampere · 2021
NVIDIA A10
The quiet default for 7B-class production inference on AWS and Oracle.
24 GB
GDDR6
0.6 TB/s
Mem bandwidth
125 TF
FP16 dense
150 W
TDP
$0.75 /hr
From · Lambda
AWS g5 instances made the A10G ubiquitous for small-model serving. 24 GB comfortably fits a 13B at 8-bit or 7B at FP16 with real batch sizes.
Where to rent a A10
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: A10 pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| Lambda | $0.75 | neocloud |
| AWS | $1.01 | g5.xlarge (A10G) |
| Modal | $1.10 | Serverless A10G |
| OCI | $2.00 | VM.GPU.A10.1 |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Ampere (2021) |
|---|---|
| Memory | 24 GB GDDR6, 0.6 TB/s |
| Interconnect | PCIe Gen4 |
| Form factor | PCIe |
| Partitioning | No MIG (time-slicing / vGPU only) |
What fits on one A10
Weights + ~2 GB runtime overhead against 22 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | 8-bit |
| Mistral 7B | 7.25B | FP16 |
| Gemma 2 9B | 9.24B | 8-bit |
| Gemma 2 27B | 27.2B | 4-bit |
| Phi-4 14B | 14.7B | 8-bit |
| gpt-oss-20b | 20.9B | 4-bit |
Best for
- 7B-13B production serving
- Graphics + AI mixed workloads
- Budget vLLM deployments