TPU v5e vs TPU v6e

TPU v6e has 2.0× the memory; TPU v6e has 2.0× the bandwidth; TPU v5e rents for 2.3× less.

The Google TPU v5e (TPU, 2023) and the Google TPU v6e (Trillium) (TPU, 2024) share the same TPU silicon generation, so this comes down to configuration and price rather than architecture. On capacity, the TPU v6e carries 2.0x the memory (32 vs 16 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the TPU v6e leads at 1.6 vs 0.82 TB/s. Raw compute favors the TPU v6e by 4.7x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding.

Price is where it settles: the TPU v5e rents from $1.20/hr against $2.70/hr for the TPU v6e, a 2.3x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the TPU v5e is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.

Google · TPU · 2023

Google TPU v5e

Google's efficiency TPU: small per chip, priced to win at pod scale.

From $1.20/hr at GCP

Google · TPU · 2024

Google TPU v6e (Trillium)

Trillium: ~4.7× v5e compute per chip. Google's mainstream training silicon.

From $2.70/hr at GCP

Head to head

■ TPU v5e   ■ TPU v6e, bars share one scale across the whole catalog.

Memory 16 GB 32 GB
Memory bandwidth 0.82 TB/s 1.6 TB/s
FP16 dense 197 TFLOPS 918 TFLOPS
FP8 dense 394 TFLOPS 1,836 TFLOPS
Power (TDP) 250 W 350 W
TPU v5eTPU v6e
Memory16 GB HBM232 GB HBM3
Bandwidth0.82 TB/s1.6 TB/s
FP16 dense197 TF918 TF
FP8 dense394 TF1,836 TF
TDP250 W350 W
InterconnectICI · 3D torus pods to 256 chipsICI · pods to 256 chips
Cheapest rental $1.20/hr $2.70/hr
$/hr per GB VRAM $7.5¢ $8.4¢

What one can run that the other can't

Only on TPU v5e

Nothing, TPU v6e runs everything TPU v5e does.

Only on TPU v6e

  • Qwen2.5 32B (4-bit)
  • Qwen2.5 Coder 32B (4-bit)
  • QwQ 32B (reasoning) (4-bit)
  • Gemma 2 27B (4-bit)
  • gpt-oss-20b (8-bit)

Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.