AM AMD Discrete GPU

AMD Instinct MI455X

Runs downloadable LLMs up to ~699B at Q4, ~2756.3-3368.8 tok/s on an 8B model.

AM
432GB · 19600 GB/s
432GB
VRAM
19600
GB/s bandwidth
699B
Max model (Q4)

What you can run

At 8k context, FP16 KV. Full-speed tok/s holds up to "max ctx"; beyond it the context spills to system RAM (CPU/PCIe offload) and drops to the degraded figure.

Model sizeFP16Q8Q4
3B 2058-2515.3 t/s
≤5950k ctx · then ~4.2-5.1 (slow)
4116-5030.7 t/s
≤5996k ctx · then ~8.4-10.3 (slow)
7350-8983.3 t/s
≤6016k ctx · then ~15-18.3 (slow)
7-8B 771.8-943.3 t/s
≤2899k ctx · then ~1.6-1.9 (slow)
1543.5-1886.5 t/s
≤2960k ctx · then ~3.2-3.9 (slow)
2756.3-3368.8 t/s
≤2987k ctx · then ~5.6-6.9 (slow)
13B 474.9-580.5 t/s
≤2258k ctx · then ~1-1.2 (slow)
949.8-1160.9 t/s
≤2337k ctx · then ~1.9-2.4 (slow)
1696.2-2073.1 t/s
≤2372k ctx · then ~3.5-4.2 (slow)
32B 192.9-235.8 t/s
≤1266k ctx · then ~0.4-0.5 (slow)
385.9-471.6 t/s
≤1388k ctx · then ~0.8-1 (slow)
689.1-842.2 t/s
≤1442k ctx · then ~1.4-1.7 (slow)
70B 88.2-107.8 t/s
≤781k ctx · then ~0.2-0.2 (slow)
176.4-215.6 t/s
≤995k ctx · then ~0.4-0.4 (slow)
315-385 t/s
≤1089k ctx · then ~0.6-0.8 (slow)
120B 51.5-62.9 t/s
≤476k ctx · then ~0.1-0.1 (slow)
102.9-125.8 t/s
≤842k ctx · then ~0.2-0.3 (slow)
183.8-224.6 t/s
≤1003k ctx · then ~0.4-0.5 (slow)
400B won't run won't run 54.4-66.5 t/s
≤323k ctx · then ~0.1-0.1 (slow)

Estimate: tok/s ≈ bandwidth ÷ (params × bytes/param) × efficiency (AMD). Cyan = fits ~397.4 GB usable. Beyond "ctx", KV spills to system RAM (CPU/PCIe offload) (rough degraded speed shown).

What else you need for a full build

Minimum supporting components for a desktop built around the AMD Instinct MI455X. Guidance — exact needs depend on your case, other parts and how much you offload to system RAM.

System RAM
≥ 512 GB · 1024 GB recommended
At least the 432 GB of VRAM; 2× lets you CPU-offload models larger than VRAM.
Motherboard
1× free PCIe x16 slot (4.0/5.0)
A full-length x16 electrical slot; PCIe 4.0 is ample, 5.0 future-proofs.
CPU
Any modern multi-core
The GPU runs inference; more cores and memory channels only matter when offloading to RAM.
Storage
≥ 1 TB NVMe SSD
Weights are large — a 70B model at Q4 is ~40 GB; budget for a few.

Run popular models on the AMD Instinct MI455X

More hardware like the AMD Instinct MI455X

About the AMD Instinct MI455X

AMD Instinct MI455X is a discrete GPU from AMD, 2026. For local genAI the numbers that matter are its 432 GB (how big a model + context fits) and 19600 GB/s bandwidth (generation speed).

Frequently asked questions

What size LLM can the AMD Instinct MI455X run?

With ~397.4 GB usable it runs up to ~699B at Q4 (~195.5B at FP16) at 8k context. Bigger spills into system RAM (CPU/PCIe offload) and slows sharply.

How many tokens/sec does the AMD Instinct MI455X generate?

Generation is bandwidth-bound. At 19600 GB/s an 8B model at Q4 runs ~2756.3-3368.8 tok/s (estimate).

Why does long context slow it down?

The KV cache for the context must also fit fast memory; once weights+KV exceed 397.4 GB it spills to system RAM (CPU/PCIe offload) and throughput collapses. Each row shows the max usable context.

What else do I need to run the AMD Instinct MI455X?

Besides the card: at least 432GB of system RAM (2× the 432GB VRAM is ideal for CPU offload), a free PCIe x16 slot, any modern multi-core CPU, and a ≥1TB NVMe SSD for model weights.