NVIDIA · Turing · 2018

NVIDIA T4

The 2018 inference stalwart. Slow by modern standards, but everywhere and nearly free.

16 GB
GDDR6
0.32 TB/s
Mem bandwidth
65 TF
FP16 dense
70 W
TDP
$0.10 /hr
From · Vast.ai

Still the cheapest CUDA-capable datacenter GPU on most clouds, and the free tier of Google Colab. Good for classic ML, small vision models, and learning CUDA.

Where to rent a T4

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: T4 pricing.

Provider $/GPU-hr vs cheapest Notes
Vast.ai $0.10 cheapest Verified-listing rate, refreshed daily
GCP $0.35 3.5× n1 + T4
AWS $0.53 5.3× g4dn.xlarge
Modal $0.59 5.9× Serverless, per-second billing

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 16 GB
Memory bandwidth 0.32 TB/s
FP16 dense 65 TFLOPS
FP8 dense -
Power (TDP) 70 W
ArchitectureTuring (2018)
Memory16 GB GDDR6, 0.32 TB/s
InterconnectPCIe Gen3
Form factorPCIe (low-profile)
PartitioningNo MIG (time-slicing / vGPU only)

What fits on one T4

Weights + ~2 GB runtime overhead against 14 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B 8-bit
Qwen2.5 7B 7.62B 8-bit
Qwen2.5 14B 14.8B 4-bit
Mistral 7B 7.25B 8-bit
Gemma 2 9B 9.24B 8-bit
Phi-4 14B 14.7B 4-bit

Best for

Compare

L4 vs T4

72 watts. The efficiency play for video, small-model inference, and high-density serving.

A10 vs T4

The quiet default for 7B-class production inference on AWS and Oracle.