NVIDIA · Turing · 2018
NVIDIA T4
The 2018 inference stalwart. Slow by modern standards, but everywhere and nearly free.
16 GB
GDDR6
0.32 TB/s
Mem bandwidth
65 TF
FP16 dense
70 W
TDP
$0.10 /hr
From · Vast.ai
Still the cheapest CUDA-capable datacenter GPU on most clouds, and the free tier of Google Colab. Good for classic ML, small vision models, and learning CUDA.
Where to rent a T4
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: T4 pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| Vast.ai | $0.10 | Verified-listing rate, refreshed daily |
| GCP | $0.35 | n1 + T4 |
| AWS | $0.53 | g4dn.xlarge |
| Modal | $0.59 | Serverless, per-second billing |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Turing (2018) |
|---|---|
| Memory | 16 GB GDDR6, 0.32 TB/s |
| Interconnect | PCIe Gen3 |
| Form factor | PCIe (low-profile) |
| Partitioning | No MIG (time-slicing / vGPU only) |
What fits on one T4
Weights + ~2 GB runtime overhead against 14 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | 8-bit |
| Qwen2.5 7B | 7.62B | 8-bit |
| Qwen2.5 14B | 14.8B | 4-bit |
| Mistral 7B | 7.25B | 8-bit |
| Gemma 2 9B | 9.24B | 8-bit |
| Phi-4 14B | 14.7B | 4-bit |
Best for
- Learning and prototyping
- Classic ML / vision inference
- Extremely low-cost batch jobs