TPU v6e vs H100 SXM

H100 SXM has 2.5× the memory; H100 SXM has 2.1× the bandwidth; H100 SXM rents for 1.4× less.

The Google TPU v6e (Trillium) (TPU, 2024) and the NVIDIA H100 SXM (Hopper, 2022) come from different vendors and different design goals, but they get cross-shopped for a reason. On capacity, the H100 SXM carries 2.5x the memory (80 vs 32 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the H100 SXM leads at 3.35 vs 1.6 TB/s.

Price is where it settles: the H100 SXM rents from $1.90/hr against $2.70/hr for the TPU v6e, a 1.4x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the H100 SXM is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

Google · TPU · 2024

Google TPU v6e (Trillium)

Trillium: ~4.7× v5e compute per chip. Google's mainstream training silicon.

From $2.70/hr at GCP

NVIDIA · Hopper · 2022

NVIDIA H100 SXM

The workhorse of the AI boom. Still the most widely available serious training and inference GPU.

From $1.90/hr at Hyperstack

Head to head

■ TPU v6e   ■ H100 SXM, bars share one scale across the whole catalog.

Memory 32 GB 80 GB
Memory bandwidth 1.6 TB/s 3.35 TB/s
FP16 dense 918 TFLOPS 990 TFLOPS
FP8 dense 1,836 TFLOPS 1,979 TFLOPS
Power (TDP) 350 W 700 W
TPU v6eH100 SXM
Memory32 GB HBM380 GB HBM3
Bandwidth1.6 TB/s3.35 TB/s
FP16 dense918 TF990 TF
FP8 dense1,836 TF1,979 TF
TDP350 W700 W
InterconnectICI · pods to 256 chipsNVLink 4 · 900 GB/s
Cheapest rental $2.70/hr $1.90/hr
$/hr per GB VRAM $8.4¢ $2.4¢

What one can run that the other can't

Only on TPU v6e

Nothing, H100 SXM runs everything TPU v6e does.

Only on H100 SXM

  • Llama 3.3 70B (4-bit)
  • Qwen2.5 72B (4-bit)
  • Mixtral 8x7B (8-bit)

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.