Ranked · 2026-09
The best GPU for inference, by the numbers
Inference cost is decided by memory bandwidth per dollar, decode streams the model's weights for every token, so bandwidth is throughput (the mechanics). We rank every rentable GPU on exactly that, with an FP8 credit for silicon that halves bytes per token, at each card's cheapest tracked rental.
Datacenter GPUs, production fleets
The tier for teams serving real traffic on infrastructure they can certify.| # | GPU | BW TB/s | From $/hr | Value score | Serves (1 GPU) |
|---|---|---|---|---|---|
| 1 | RTX PRO 6000FP8 | 1.79 | $0.50 | 5.7 | 70B-class single GPU |
| 2 | GH200FP8 | 4 | $1.35 | 4.7 | 70B-class single GPU |
| 3 | MI300XFP8 | 5.3 | $1.85 | 4.6 | 70B-class single GPU |
| 4 | V100 | 0.9 | $0.21 | 4.3 | up to ~13B (4-bit: 32B) |
| 5 | MI325XFP8 | 6 | $2.40 | 4.0 | 70B-class single GPU |
| 6 | B200FP8 | 8 | $3.75 | 3.4 | 70B-class single GPU |
| 7 | H200FP8 | 4.8 | $2.30 | 3.3 | 70B-class single GPU |
| 8 | A100 40GB | 1.56 | $0.48 | 3.3 | up to ~34B (4-bit) |
| 9 | T4 | 0.32 | $0.10 | 3.2 | small models |
| 10 | H100 SXMFP8 | 3.35 | $1.90 | 2.8 | 70B-class single GPU |
| 11 | Gaudi 3FP8 | 3.7 | $2.30 | 2.6 | 70B-class single GPU |
| 12 | A100 80GB | 2 | $0.87 | 2.3 | 70B-class single GPU |
| 13 | H100 PCIeFP8 | 2 | $1.60 | 2.0 | 70B-class single GPU |
| 14 | L40SFP8 | 0.864 | $0.80 | 1.7 | up to ~34B (4-bit) |
| 15 | L4FP8 | 0.3 | $0.32 | 1.5 | up to ~13B (4-bit: 32B) |
| 16 | A10 | 0.6 | $0.75 | 0.8 | up to ~13B (4-bit: 32B) |
Consumer cards, marketplace rentals
Unbeatable value scores, with marketplace trust and reliability trade-offs priced in.Value score = TB/s ÷ cheapest $/hr, ×1.6 where FP8 is supported. Cheapest tracked on-demand prices, 2026-09 (methodology).
How to actually choose
- First, VRAM decides eligibility, a great value score means nothing if the model doesn't fit. Run yours through the VRAM calculator.
- Serving a 70B-class model cheaply: A100 80GB (4-bit) for throughput on a budget, H100 when FP8 and latency matter, the decision guide.
- Serving 7B-13B: consumer cards (4090, 3090) dominate on price where marketplace hosting is acceptable; L4/A10 for boring reliability.
- Long context or 100B+: memory capacity becomes the constraint, H200, MI300X, MI325X.
- Owning instead of renting: used A100s at ~$6,800 turn into ~$0.35/hr at healthy utilization, the ROI calculator and the second-life story cover it.
Why trust this ranking: it's a formula over public specs and tracked prices, not editorial preference, and the formula is printed above so you can disagree with it precisely. When prices move, the ranking moves.