The GPU database

21 accelerators that matter for AI, with the numbers that decide what they can run, and the cheapest on-demand price across the clouds we track.

Datacenter GPUs

GPU VRAM BW TB/s FP16 TF From $/hr Arch
MI325X 256 GB 6 1,307 $2.40 CDNA 3 '24
B200 192 GB 8 2,250 $3.75 Blackwell '24
MI300X 192 GB 5.3 1,307 $1.85 CDNA 3 '23
H200 141 GB 4.8 990 $2.30 Hopper '24
Gaudi 3 128 GB 3.7 918 $2.30 Gaudi '24
GH200 96 GB 4 990 $1.35 Grace Hopper '23
RTX PRO 6000 96 GB 1.79 250 $0.50 Blackwell '25
H100 SXM 80 GB 3.35 990 $1.90 Hopper '22
H100 PCIe 80 GB 2 756 $1.60 Hopper '22
A100 80GB 80 GB 2 312 $0.87 Ampere '20
L40S 48 GB 0.864 362 $0.80 Ada Lovelace '23
A100 40GB 40 GB 1.56 312 $0.48 Ampere '20
V100 32 GB 0.9 125 $0.21 Volta '17
A10 24 GB 0.6 125 $0.75 Ampere '21
L4 24 GB 0.3 121 $0.32 Ada Lovelace '23
T4 16 GB 0.32 65 $0.10 Turing '18

Consumer cards (marketplace rentals)

GPU VRAM BW TB/s FP16 TF From $/hr Arch
RTX 5090 32 GB 1.79 210 $0.54 Blackwell '25
RTX 4090 24 GB 1.01 165 $0.34 Ada Lovelace '22
RTX 3090 24 GB 0.936 71 $0.18 Ampere '20

AI ASICs

GPU VRAM BW TB/s FP16 TF From $/hr Arch
TPU v6e 32 GB 1.6 918 $2.70 TPU '24
TPU v5e 16 GB 0.82 197 $1.20 TPU '23

Reading the table: for LLM inference, memory bandwidth (TB/s) predicts tokens/second better than FLOPS, decoding is memory-bound. FLOPS matter for training, prefill, and diffusion models. VRAM decides what fits at all: see the VRAM math.