TPU v5e vs TPU v6e
TPU v6e has 2.0× the memory; TPU v6e has 2.0× the bandwidth; TPU v5e rents for 2.3× less.
The Google TPU v5e (TPU, 2023) and the Google TPU v6e (Trillium) (TPU, 2024) share the same TPU silicon generation, so this comes down to configuration and price rather than architecture. On capacity, the TPU v6e carries 2.0x the memory (32 vs 16 GB), which sets what each can hold at all. On memory bandwidth, the number that governs LLM serving speed, the TPU v6e leads at 1.6 vs 0.82 TB/s. Raw compute favors the TPU v6e by 4.7x, which matters for training, prefill-heavy traffic, and diffusion work more than for chat-style decoding.
Price is where it settles: the TPU v5e rents from $1.20/hr against $2.70/hr for the TPU v6e, a 2.3x gap that the performance numbers above only partly close. Per gigabyte of VRAM, the TPU v5e is the cheaper rental, worth knowing if your model is capacity-bound rather than speed-bound. For how these numbers translate into tokens per dollar, the inference cost estimator runs both cards against any model. Everything below is the underlying data.
Google · TPU · 2023
Google TPU v5e
Google's efficiency TPU: small per chip, priced to win at pod scale.
From $1.20/hr at GCP
Google · TPU · 2024
Google TPU v6e (Trillium)
Trillium: ~4.7× v5e compute per chip. Google's mainstream training silicon.
From $2.70/hr at GCP
Head to head
■ TPU v5e ■ TPU v6e, bars share one scale across the whole catalog.
| TPU v5e | TPU v6e | |
|---|---|---|
| Memory | 16 GB HBM2 | 32 GB HBM3 |
| Bandwidth | 0.82 TB/s | 1.6 TB/s |
| FP16 dense | 197 TF | 918 TF |
| FP8 dense | 394 TF | 1,836 TF |
| TDP | 250 W | 350 W |
| Interconnect | ICI · 3D torus pods to 256 chips | ICI · pods to 256 chips |
| Cheapest rental | $1.20/hr | $2.70/hr |
| $/hr per GB VRAM | $7.5¢ | $8.4¢ |
What one can run that the other can't
Only on TPU v5e
Nothing, TPU v6e runs everything TPU v5e does.
Only on TPU v6e
- Qwen2.5 32B (4-bit)
- Qwen2.5 Coder 32B (4-bit)
- QwQ 32B (reasoning) (4-bit)
- Gemma 2 27B (4-bit)
- gpt-oss-20b (8-bit)
Rule of thumb: for LLM serving, prefer the GPU with more memory bandwidth per dollar; for training and prefill-heavy work, prefer FLOPS per dollar; and if the model doesn't fit in VRAM, none of the other numbers matter. Sanity-check with the VRAM calculator.