Google · TPU · 2024
Google TPU v6e (Trillium)
Trillium: ~4.7× v5e compute per chip. Google's mainstream training silicon.
32 GB
HBM3
1.6 TB/s
Mem bandwidth
918 TF
FP16 dense
1,836 TF
FP8 dense
350 W
TDP
$2.70 /hr
From · GCP
The TPU generation Google trains Gemini on (alongside v5p). Rentable only on Google Cloud; the tight JAX/XLA integration is the draw and the lock-in.
Where to rent it
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: TPU v6e pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| GCP | $2.70 | per chip-hour, us-east5 |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | TPU (2024) |
|---|---|
| Memory | 32 GB HBM3, 1.6 TB/s |
| Interconnect | ICI · pods to 256 chips |
| Form factor | Cloud pod (GCP only) |
| Partitioning | No MIG (time-slicing / vGPU only) |
What fits on one TPU v6e
Weights + ~2 GB runtime overhead against 29 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | 8-bit |
| Qwen2.5 32B | 32.8B | 4-bit |
| Qwen2.5 Coder 32B | 32.8B | 4-bit |
| QwQ 32B (reasoning) | 32.8B | 4-bit |
| Mistral 7B | 7.25B | FP16 |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | 4-bit |
| Phi-4 14B | 14.7B | 8-bit |
| gpt-oss-20b | 20.9B | 8-bit |
Best for
- Large-scale JAX/XLA training
- Gemma / Gemini-adjacent stacks
- Teams already on GCP