New: connect Claude & other AIs to GenAIList over MCP — research the catalog and contribute to the shared knowledge base. Learn how →
AMD Instinct MI455X
Runs downloadable LLMs up to ~699B at Q4, ~2756.3-3368.8 tok/s on an 8B model.
What you can run
At 8k context, FP16 KV. Full-speed tok/s holds up to "max ctx"; beyond it the context spills to system RAM (CPU/PCIe offload) and drops to the degraded figure.
| Model size | FP16 | Q8 | Q4 |
|---|---|---|---|
| 3B |
2058-2515.3 t/s ≤5950k ctx · then ~4.2-5.1 (slow) |
4116-5030.7 t/s ≤5996k ctx · then ~8.4-10.3 (slow) |
7350-8983.3 t/s ≤6016k ctx · then ~15-18.3 (slow) |
| 7-8B |
771.8-943.3 t/s ≤2899k ctx · then ~1.6-1.9 (slow) |
1543.5-1886.5 t/s ≤2960k ctx · then ~3.2-3.9 (slow) |
2756.3-3368.8 t/s ≤2987k ctx · then ~5.6-6.9 (slow) |
| 13B |
474.9-580.5 t/s ≤2258k ctx · then ~1-1.2 (slow) |
949.8-1160.9 t/s ≤2337k ctx · then ~1.9-2.4 (slow) |
1696.2-2073.1 t/s ≤2372k ctx · then ~3.5-4.2 (slow) |
| 32B |
192.9-235.8 t/s ≤1266k ctx · then ~0.4-0.5 (slow) |
385.9-471.6 t/s ≤1388k ctx · then ~0.8-1 (slow) |
689.1-842.2 t/s ≤1442k ctx · then ~1.4-1.7 (slow) |
| 70B |
88.2-107.8 t/s ≤781k ctx · then ~0.2-0.2 (slow) |
176.4-215.6 t/s ≤995k ctx · then ~0.4-0.4 (slow) |
315-385 t/s ≤1089k ctx · then ~0.6-0.8 (slow) |
| 120B |
51.5-62.9 t/s ≤476k ctx · then ~0.1-0.1 (slow) |
102.9-125.8 t/s ≤842k ctx · then ~0.2-0.3 (slow) |
183.8-224.6 t/s ≤1003k ctx · then ~0.4-0.5 (slow) |
| 400B | won't run | won't run |
54.4-66.5 t/s ≤323k ctx · then ~0.1-0.1 (slow) |
Estimate: tok/s ≈ bandwidth ÷ (params × bytes/param) × efficiency (AMD). Cyan = fits ~397.4 GB usable. Beyond "ctx", KV spills to system RAM (CPU/PCIe offload) (rough degraded speed shown).
What else you need for a full build▶
Minimum supporting components for a desktop built around the AMD Instinct MI455X. Guidance — exact needs depend on your case, other parts and how much you offload to system RAM.
Run popular models on the AMD Instinct MI455X
More hardware like the AMD Instinct MI455X
About the AMD Instinct MI455X
AMD Instinct MI455X is a discrete GPU from AMD, 2026. For local genAI the numbers that matter are its 432 GB (how big a model + context fits) and 19600 GB/s bandwidth (generation speed).
Frequently asked questions
What size LLM can the AMD Instinct MI455X run?▶
With ~397.4 GB usable it runs up to ~699B at Q4 (~195.5B at FP16) at 8k context. Bigger spills into system RAM (CPU/PCIe offload) and slows sharply.
How many tokens/sec does the AMD Instinct MI455X generate?▶
Generation is bandwidth-bound. At 19600 GB/s an 8B model at Q4 runs ~2756.3-3368.8 tok/s (estimate).
Why does long context slow it down?▶
The KV cache for the context must also fit fast memory; once weights+KV exceed 397.4 GB it spills to system RAM (CPU/PCIe offload) and throughput collapses. Each row shows the max usable context.
What else do I need to run the AMD Instinct MI455X?▶
Besides the card: at least 432GB of system RAM (2× the 432GB VRAM is ideal for CPU offload), a free PCIe x16 slot, any modern multi-core CPU, and a ≥1TB NVMe SSD for model weights.