A100 vs H100 for inference
7 min read · updated 2026-09
The H100 is unambiguously the better GPU. It is not unambiguously the better deal, for inference, the answer flips with your workload shape. Here's the line, drawn with numbers.
The raw gap
| A100 80GB | H100 SXM | Gap | |
|---|---|---|---|
| Memory bandwidth | 2.0 TB/s | 3.35 TB/s | 1.7× |
| FP16 compute | 312 TF | 990 TF | 3.2× |
| FP8 compute | - | 1,979 TF | H100 only |
| VRAM | 80 GB | 80 GB | even |
| Cheapest rent (2026-09) | $0.85/hr | $1.65/hr | 1.9× |
| Typical neocloud rent | $1.29/hr | $2.49/hr | 1.9× |
| Used unit price | ~$6,800 | ~$23,000 | 3.4× |
Decode throughput tracks bandwidth (why), so at FP16/INT8 the H100 produces ~1.7× the tokens for ~1.9× the price: the A100 makes slightly cheaper tokens. The H100's case rests on what the A100 can't do at all.
Where the H100 genuinely wins
- FP8 serving. Hopper's Transformer Engine runs weights and KV cache in FP8, halving bytes per token. Now it's ~3.4× A100 throughput for ~1.9× price, the H100 makes cheaper tokens. This is the single biggest decider, and it applies when your model has a good FP8 build (mainstream architectures do).
- Latency SLOs. If your product needs <2s first token and fast streaming on a 70B at real concurrency, the A100 forces you to shrink batches, its cost advantage evaporates.
- Prefill-heavy traffic. RAG with 20K-token prompts is compute-bound during prefill; 3.2× FLOPS matters there.
- Rack density. Same power budget, ~2-3× tokens per rack. If you're space- or power-constrained, buy the newer card.
Where the A100 wins
- Throughput-first serving at INT8/4-bit, batch summarization, evals, embeddings, agent fleets, overnight pipelines: workloads where tokens/dollar beats tokens/second.
- 7B-14B production endpoints, an H100 is wasted on a 8B model; an A100 (or its MIG slices) is right-sized.
- Fine-tuning economics, QLoRA on a 70B doesn't need FP8; it needs 80 GB and patience at $0.85/hr.
- Buying hardware outright, at ~$6,800 vs ~$23,000 used, the A100 breaks even against rentals in months, not years. The ROI calculator does your exact numbers.
The one-question shortcut
"Will I serve this model in FP8, and do I care about latency?" Two yeses: H100. Two nos: A100 and bank the difference. One yes: run both through the inference cost estimator, it models exactly this trade with current prices.