// RUN โ€” FITS

Can the Jetson Thor T2000 run Llama 3.2 11B?

Yes โ€” here's how fast and up to what context, with real measured numbers.

Llama 3.2 11B on the Jetson Thor T2000

QuantFits?Max contextSpeedBeyond context
FP16 no โ€” โ€” โ€”
Q8 yes 14k 9.4-11.2 t/s ~0.7-0.8 (slow)
Q4 yes 43k 16.8-20.1 t/s ~1.2-1.5 (slow)

Estimate (memory-bound). Beyond "max context" the KV cache spills to SSD swap and speed drops to the "beyond" figure.

FAQ

Can the Jetson Thor T2000 run Llama 3.2 11B?โ–ถ

Yes. At Q4 it generates ~16.8-20.1 tok/s and fits up to 43k context; beyond that it spills to SSD swap and slows to ~1.2-1.5 tok/s.

How much context fits?โ–ถ

At Q4, up to about 43k tokens stay in fast memory; longer context spills to SSD swap and throughput collapses.