Intel · Gaudi · 2024
Intel Gaudi 3
Intel's price-performance play, with Ethernet-native scale-out instead of proprietary interconnect.
Integrated 200GbE means clusters use standard Ethernet switching. Software (Synapse AI, PyTorch via Habana) covers mainstream transformer work; expect friction off the beaten path. IBM Cloud and Intel Tiber offer it aggressively priced.
Deep dive: Gaudi 3, the accelerator nobody is fighting over →
Where to rent a Gaudi 3
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: Gaudi 3 pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| FluidStack | $2.30 | Via IBM Cloud capacity |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Gaudi (2024) |
|---|---|
| Memory | 128 GB HBM2e, 3.7 TB/s |
| Interconnect | 24×200GbE RoCE on-chip |
| Form factor | OAM |
| Partitioning | No MIG (time-slicing / vGPU only) |
What fits on one Gaudi 3
Weights + ~2 GB runtime overhead against 115 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Llama 3.3 70B | 70.6B | 8-bit |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | FP16 |
| Qwen2.5 32B | 32.8B | FP16 |
| Qwen2.5 Coder 32B | 32.8B | FP16 |
| Qwen2.5 72B | 72.7B | 8-bit |
| QwQ 32B (reasoning) | 32.8B | FP16 |
| Mistral 7B | 7.25B | FP16 |
| Mixtral 8x7B | 46.7B | FP16 |
| Mixtral 8x22B | 140.6B | 4-bit |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | FP16 |
| Phi-4 14B | 14.7B | FP16 |
| gpt-oss-20b | 20.9B | FP16 |
| gpt-oss-120b | 116.8B | 4-bit |
Best for
- Cost-focused training
- Ethernet-standard cluster builds
- Mainstream transformer workloads
Compare
H100 SXM vs Gaudi 3
The workhorse of the AI boom. Still the most widely available serious training and inference GPU.
H200 vs Gaudi 3
An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.
Gaudi 3 vs MI300X
AMD's answer to Hopper: 192 GB on one GPU, a 70B model in FP16 fits with room to spare.
Gaudi 3 vs A100 80GB
The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.