AMD · CDNA 3 · 2023
AMD Instinct MI300X
AMD's answer to Hopper: 192 GB on one GPU, a 70B model in FP16 fits with room to spare.
More memory and bandwidth than an H200 on paper. Software is the real question: ROCm plus vLLM is production-grade for mainstream architectures, but the CUDA ecosystem's long tail is still NVIDIA's moat. Azure, Oracle and several neoclouds offer it, often undercutting H100 pricing per GB.
Where to rent a MI300X
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: MI300X pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| Vast.ai | $1.85 | Limited listings |
| RunPod | $2.39 | neocloud |
| Crusoe | $2.60 | neocloud |
| DigitalOcean | $3.06 | neocloud |
| OCI | $3.75 | BM.GPU.MI300X.8 ÷ 8 |
| Azure | $5.42 | ND MI300X v5 ÷ 8 |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | CDNA 3 (2023) |
|---|---|
| Memory | 192 GB HBM3, 5.3 TB/s |
| Interconnect | Infinity Fabric · 896 GB/s |
| Form factor | OAM |
| Partitioning | No MIG (time-slicing / vGPU only) |
What fits on one MI300X
Weights + ~2 GB runtime overhead against 173 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Llama 3.3 70B | 70.6B | FP16 |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | FP16 |
| Qwen2.5 32B | 32.8B | FP16 |
| Qwen2.5 Coder 32B | 32.8B | FP16 |
| Qwen2.5 72B | 72.7B | FP16 |
| QwQ 32B (reasoning) | 32.8B | FP16 |
| Mistral 7B | 7.25B | FP16 |
| Mixtral 8x7B | 46.7B | FP16 |
| Mixtral 8x22B | 140.6B | 8-bit |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | FP16 |
| Phi-4 14B | 14.7B | FP16 |
| gpt-oss-20b | 20.9B | FP16 |
| gpt-oss-120b | 116.8B | 8-bit |
Best for
- Single-GPU 70B FP16 inference
- Memory-bound serving at scale
- Escaping the NVIDIA queue
Compare
H100 SXM vs MI300X
The workhorse of the AI boom. Still the most widely available serious training and inference GPU.
H200 vs MI300X
An H100 with 76% more memory and 43% more bandwidth, the practical choice for large-model inference.
MI300X vs MI325X
256 GB of HBM3e, the largest memory pool on any single accelerator you can rent.
Gaudi 3 vs MI300X
Intel's price-performance play, with Ethernet-native scale-out instead of proprietary interconnect.