Google · TPU · 2024

Google TPU v6e (Trillium)

Trillium: ~4.7× v5e compute per chip. Google's mainstream training silicon.

32 GB
HBM3
1.6 TB/s
Mem bandwidth
918 TF
FP16 dense
1,836 TF
FP8 dense
350 W
TDP
$2.70 /hr
From · GCP

The TPU generation Google trains Gemini on (alongside v5p). Rentable only on Google Cloud; the tight JAX/XLA integration is the draw and the lock-in.

Where to rent it

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: TPU v6e pricing.

Provider $/GPU-hr vs cheapest Notes
GCP $2.70 cheapest per chip-hour, us-east5

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 32 GB
Memory bandwidth 1.6 TB/s
FP16 dense 918 TFLOPS
FP8 dense 1,836 TFLOPS
Power (TDP) 350 W
ArchitectureTPU (2024)
Memory32 GB HBM3, 1.6 TB/s
InterconnectICI · pods to 256 chips
Form factorCloud pod (GCP only)
PartitioningNo MIG (time-slicing / vGPU only)

What fits on one TPU v6e

Weights + ~2 GB runtime overhead against 29 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B 8-bit
Qwen2.5 32B 32.8B 4-bit
Qwen2.5 Coder 32B 32.8B 4-bit
QwQ 32B (reasoning) 32.8B 4-bit
Mistral 7B 7.25B FP16
Gemma 2 9B 9.24B FP16
Gemma 2 27B 27.2B 4-bit
Phi-4 14B 14.7B 8-bit
gpt-oss-20b 20.9B 8-bit

Best for

Compare

TPU v5e vs TPU v6e

Google's efficiency TPU: small per chip, priced to win at pod scale.

TPU v6e vs H100 SXM

The workhorse of the AI boom. Still the most widely available serious training and inference GPU.