Gaudi 3: the accelerator nobody is fighting over
9 min read · updated 2026-09
Every other accelerator on this site is priced by demand. Gaudi 3 is priced by the lack of it. Intel built a genuinely competitive inference chip, failed to sell it, and cancelled its successor's commercial run. The market read that as a verdict on the silicon. It is not. It is a discount.
What Intel actually built
Strip the logo off the spec sheet and Gaudi 3 reads like a card the market should be fighting over:
| Gaudi 3 | H100 SXM | A100 80GB | |
|---|---|---|---|
| Memory | 128 GB HBM2e | 80 GB HBM3 | 80 GB HBM2e |
| Bandwidth | 3.7 TB/s | 3.35 TB/s | 2.0 TB/s |
| BF16 dense | 918 TF | 990 TF | 312 TF |
| FP8 dense | 1,835 TF | 1,979 TF | unsupported |
| Scale-out | 24x 200GbE on-chip | NVLink + InfiniBand | NVLink + InfiniBand |
| Tracked rental from | $2.30/hr | $1.65/hr | $0.85/hr |
More memory than an H100, more bandwidth than an H100, 93% of its FP8 compute, and, since LLM decoding is bound by memory bandwidth, a spec profile aimed squarely at inference. A 4-bit 70B model fits on one Gaudi 3 with generous KV cache headroom. A FP8 70B fits with room to spare. On our own bandwidth-per-dollar ranking logic, Gaudi 3 at its going rate lands in the top tier of the datacenter class.
The failure, plainly
None of that sold chips. Intel set a modest revenue target for Gaudi in 2024, roughly one percent of NVIDIA's datacenter quarter, and publicly missed it. In early 2025 it cancelled the commercial launch of Falcon Shores, Gaudi's successor, keeping it as an internal test vehicle and pointing customers at a future part. The buyers with the biggest checkbooks concluded what they always conclude about a roadmap gap: nobody gets fired for buying NVIDIA.
Here is the thing about a failed accelerator, though: failure is a demand event, not a supply event. The chips that exist still run at full speed. The software that worked yesterday still works. What changed is the price, and only the price. This is the same mechanism that made the A100 the value king, run faster and harder. The A100 got cheap through five years of orderly depreciation. Gaudi 3 got cheap in about eighteen months, through reputational collapse.
Where the discount is real
- Mainstream transformer serving. Llama, Qwen, Mistral and friends run well through vLLM's Gaudi backend and Intel's PyTorch stack. If your production workload is "serve a popular open model efficiently," the CUDA moat barely touches you, and you are the buyer this discount was accidentally built for.
- Ethernet-native clusters. Gaudi 3 scales out over standard 200GbE RoCE, no InfiniBand fabric, no NVLink switch tax, no proprietary networking vendor. For a cost-driven cluster, the networking line item drops by real money, and your network engineers already know how to run it.
- Throughput economics. The buyer profile is the same one described in cost per token: batch and background inference, fine-tuning of mainstream architectures, token factories. Latency-obsessed products and research groups chasing bleeding-edge techniques should stay on NVIDIA, and it is fine to say so out loud.
Where the discount is earned
The price is low for reasons, and honesty about them is what makes the bargain real rather than naive:
- The software long tail. Anything CUDA-only, custom kernels, niche quantization schemes, day-one support for new architectures, arrives on Gaudi late or never. Budget engineering time for the first deployment, and prototype your exact model before committing a fleet.
- Roadmap risk. There is no commercial Gaudi 4. Whatever Intel ships next targets a different architecture. Buy Gaudi 3 for what it does today at today's price, amortized over a duty cycle you control, never for a future upgrade path.
- Resale value near zero. An H100 holds value because everyone wants one. Gaudi 3's exit price should be modeled at scrap. That pushes the math toward renting it, or toward contract pricing where someone else carries the residual.
- Thin availability. A handful of clouds carry it (IBM Cloud most visibly). Fewer landlords means less price competition on rentals than raw demand would suggest.
The verdict
Gaudi 3 is what a value market looks like before consensus arrives: strong silicon, wounded brand, price set by fear rather than physics. Treat it like the special situation it is. Model the workload in the VRAM calculator, price it against A100 and H100 options in the cost estimator, pilot on rented capacity before buying anything, and hold NVIDIA in the slots where the ecosystem premium earns its keep. The market is efficient about hardware it loves. It is sloppy about hardware it has written off, and sloppy markets are where the cost-per-token wins live.