Intel · Gaudi · 2024

Intel Gaudi 3

Intel's price-performance play, with Ethernet-native scale-out instead of proprietary interconnect.

128 GB
HBM2e
3.7 TB/s
Mem bandwidth
918 TF
FP16 dense
1,835 TF
FP8 dense
900 W
TDP
$2.30 /hr
From · FluidStack

Integrated 200GbE means clusters use standard Ethernet switching. Software (Synapse AI, PyTorch via Habana) covers mainstream transformer work; expect friction off the beaten path. IBM Cloud and Intel Tiber offer it aggressively priced.

Deep dive: Gaudi 3, the accelerator nobody is fighting over →

Where to rent a Gaudi 3

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: Gaudi 3 pricing.

Provider $/GPU-hr vs cheapest Notes
FluidStack $2.30 cheapest Via IBM Cloud capacity

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 128 GB
Memory bandwidth 3.7 TB/s
FP16 dense 918 TFLOPS
FP8 dense 1,835 TFLOPS
Power (TDP) 900 W
ArchitectureGaudi (2024)
Memory128 GB HBM2e, 3.7 TB/s
Interconnect24×200GbE RoCE on-chip
Form factorOAM
PartitioningNo MIG (time-slicing / vGPU only)

What fits on one Gaudi 3

Weights + ~2 GB runtime overhead against 115 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Llama 3.3 70B 70.6B 8-bit
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B FP16
Qwen2.5 32B 32.8B FP16
Qwen2.5 Coder 32B 32.8B FP16
Qwen2.5 72B 72.7B 8-bit
QwQ 32B (reasoning) 32.8B FP16
Mistral 7B 7.25B FP16
Mixtral 8x7B 46.7B FP16
Mixtral 8x22B 140.6B 4-bit
Gemma 2 9B 9.24B FP16
Gemma 2 27B 27.2B FP16
Phi-4 14B 14.7B FP16
gpt-oss-20b 20.9B FP16
gpt-oss-120b 116.8B 4-bit

Best for

Compare

H100 SXM vs Gaudi 3

The workhorse of the AI boom. Still the most widely available serious training and inference GPU.

H200 vs Gaudi 3

An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.

Gaudi 3 vs MI300X

AMD's answer to Hopper: 192 GB on one GPU, a 70B model in FP16 fits with room to spare.

Gaudi 3 vs A100 80GB

The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.