NVIDIA · Volta · 2017

NVIDIA V100

The first Tensor Core GPU. Retired from the frontier, still cheap and capable for small models.

32 GB
HBM2
0.9 TB/s
Mem bandwidth
125 TF
FP16 dense
300 W
TDP
$0.21 /hr
From · Vast.ai

No BF16 support complicates modern training recipes. As a cheap rental it remains serviceable for inference of models up to ~13B quantized.

Where to rent a V100

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: V100 pricing.

Provider $/GPU-hr vs cheapest Notes
Vast.ai $0.21 cheapest Verified-listing rate, refreshed daily
RunPod $0.29 1.4× neocloud
OVHcloud $0.88 4.2× hyperscaler
AWS $3.06 14.6× p3.2xlarge, legacy pricing

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 32 GB
Memory bandwidth 0.9 TB/s
FP16 dense 125 TFLOPS
FP8 dense -
Power (TDP) 300 W
ArchitectureVolta (2017)
Memory32 GB HBM2, 0.9 TB/s
InterconnectNVLink 2 · 300 GB/s
Form factorSXM / PCIe
PartitioningNo MIG (time-slicing / vGPU only)

What fits on one V100

Weights + ~2 GB runtime overhead against 29 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B 8-bit
Qwen2.5 32B 32.8B 4-bit
Qwen2.5 Coder 32B 32.8B 4-bit
QwQ 32B (reasoning) 32.8B 4-bit
Mistral 7B 7.25B FP16
Gemma 2 9B 9.24B FP16
Gemma 2 27B 27.2B 4-bit
Phi-4 14B 14.7B 8-bit
gpt-oss-20b 20.9B 8-bit

Best for

Compare

A100 80GB vs V100

The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.