Ranked · 2026-09

The best GPU for inference, by the numbers

Inference cost is decided by memory bandwidth per dollar, decode streams the model's weights for every token, so bandwidth is throughput (the mechanics). We rank every rentable GPU on exactly that, with an FP8 credit for silicon that halves bytes per token, at each card's cheapest tracked rental.

Datacenter GPUs, production fleets

The tier for teams serving real traffic on infrastructure they can certify.
#GPUBW TB/sFrom $/hrValue scoreServes (1 GPU)
1 RTX PRO 6000FP8 1.79 $0.50 5.7 70B-class single GPU
2 GH200FP8 4 $1.35 4.7 70B-class single GPU
3 MI300XFP8 5.3 $1.85 4.6 70B-class single GPU
4 V100 0.9 $0.21 4.3 up to ~13B (4-bit: 32B)
5 MI325XFP8 6 $2.40 4.0 70B-class single GPU
6 B200FP8 8 $3.75 3.4 70B-class single GPU
7 H200FP8 4.8 $2.30 3.3 70B-class single GPU
8 A100 40GB 1.56 $0.48 3.3 up to ~34B (4-bit)
9 T4 0.32 $0.10 3.2 small models
10 H100 SXMFP8 3.35 $1.90 2.8 70B-class single GPU
11 Gaudi 3FP8 3.7 $2.30 2.6 70B-class single GPU
12 A100 80GB 2 $0.87 2.3 70B-class single GPU
13 H100 PCIeFP8 2 $1.60 2.0 70B-class single GPU
14 L40SFP8 0.864 $0.80 1.7 up to ~34B (4-bit)
15 L4FP8 0.3 $0.32 1.5 up to ~13B (4-bit: 32B)
16 A10 0.6 $0.75 0.8 up to ~13B (4-bit: 32B)

Consumer cards, marketplace rentals

Unbeatable value scores, with marketplace trust and reliability trade-offs priced in.
#GPUBW TB/sFrom $/hrValue scoreServes (1 GPU)
1 RTX 5090FP8 1.79 $0.54 5.3 up to ~13B (4-bit: 32B)
2 RTX 3090 0.936 $0.18 5.2 up to ~13B (4-bit: 32B)
3 RTX 4090FP8 1.01 $0.34 4.8 up to ~13B (4-bit: 32B)

Value score = TB/s ÷ cheapest $/hr, ×1.6 where FP8 is supported. Cheapest tracked on-demand prices, 2026-09 (methodology).

How to actually choose

Why trust this ranking: it's a formula over public specs and tracked prices, not editorial preference, and the formula is printed above so you can disagree with it precisely. When prices move, the ranking moves.