Google · TPU · 2023
Google TPU v5e
Google's efficiency TPU: small per chip, priced to win at pod scale.
16 GB
HBM2
0.82 TB/s
Mem bandwidth
197 TF
FP16 dense
394 TF
FP8 dense
250 W
TDP
$1.20 /hr
From · GCP
16 GB per chip means sharding is mandatory for big models, but pod pricing per effective FLOP is among the lowest anywhere. JAX is first-class; PyTorch/XLA works with caveats.
Where to rent it
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: TPU v5e pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| GCP | $1.20 | per chip-hour, us-west4 |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | TPU (2023) |
|---|---|
| Memory | 16 GB HBM2, 0.82 TB/s |
| Interconnect | ICI · 3D torus pods to 256 chips |
| Form factor | Cloud pod (GCP only) |
| Partitioning | No MIG (time-slicing / vGPU only) |
What fits on one TPU v5e
Weights + ~2 GB runtime overhead against 14 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | 8-bit |
| Qwen2.5 7B | 7.62B | 8-bit |
| Qwen2.5 14B | 14.8B | 4-bit |
| Mistral 7B | 7.25B | 8-bit |
| Gemma 2 9B | 9.24B | 8-bit |
| Phi-4 14B | 14.7B | 4-bit |
Best for
- JAX training at pod scale
- Cost-optimized inference of sharded models
- Embedding / rerank fleets