// RUN โ€” FITS

Can the AMD Instinct MI455X run Llama Nemotron Super 49B?

Yes โ€” here's how fast and up to what context, with real measured numbers.

Llama Nemotron Super 49B on the AMD Instinct MI455X

QuantFits?Max contextSpeedBeyond context
FP16 yes 909k 126-154 t/s ~0.3-0.3 (slow)
Q8 yes 1059k 252-308 t/s ~0.5-0.6 (slow)
Q4 yes 1125k 450-550 t/s ~0.9-1.1 (slow)

Estimate (memory-bound). Beyond "max context" the KV cache spills to system RAM offload (depends on your system RAM) and speed drops to the "beyond" figure.

FAQ

Can the AMD Instinct MI455X run Llama Nemotron Super 49B?โ–ถ

Yes. At Q4 it generates ~450-550 tok/s and fits up to 1125k context; beyond that it spills to system RAM offload (depends on your system RAM) and slows to ~0.9-1.1 tok/s.

How much context fits?โ–ถ

At Q4, up to about 1125k tokens stay in fast memory; longer context spills to system RAM offload (depends on your system RAM) and throughput collapses.