NV NVIDIA Unified memory

Jetson Thor T3000

Runs downloadable LLMs up to ~43.5B at Q4, ~44.5-53 tok/s on an 8B model.

NV
32GB · 273 GB/s
32GB
Unified memory
273
GB/s bandwidth
43.5B
Max model (Q4)

What you can run

At 8k context, FP16 KV. Full-speed tok/s holds up to "max ctx"; beyond it the context spills to SSD swap and drops to the degraded figure.

Model sizeFP16Q8Q4
3B 33.2-39.6 t/s
≤325k ctx · then ~1.2-1.5 (slow)
66.4-79.2 t/s
≤371k ctx · then ~2.4-2.9 (slow)
118.6-141.4 t/s
≤391k ctx · then ~4.3-5.2 (slow)
7-8B 12.5-14.8 t/s
≤86k ctx · then ~0.5-0.5 (slow)
24.9-29.7 t/s
≤147k ctx · then ~0.9-1.1 (slow)
44.5-53 t/s
≤174k ctx · then ~1.6-1.9 (slow)
13B 7.7-9.1 t/s
≤8k ctx · then ~0.3-0.3 (slow)
15.3-18.3 t/s
≤87k ctx · then ~0.6-0.7 (slow)
27.4-32.6 t/s
≤122k ctx · then ~1-1.2 (slow)
32B won't run won't run 11.1-13.3 t/s
≤36k ctx · then ~0.4-0.5 (slow)
70B won't run won't run won't run
120B won't run won't run won't run
400B won't run won't run won't run

Estimate: tok/s ≈ bandwidth ÷ (params × bytes/param) × efficiency (NVIDIA). Cyan = fits ~28.8 GB usable. Beyond "ctx", KV spills to SSD swap (rough degraded speed shown).

Run popular models on the Jetson Thor T3000

More hardware like the Jetson Thor T3000

About the Jetson Thor T3000

Jetson Thor T3000 is a unified-memory system from NVIDIA, 2027. For local genAI the numbers that matter are its 32 GB (how big a model + context fits) and 273 GB/s bandwidth (generation speed).

Frequently asked questions

What size LLM can the Jetson Thor T3000 run?â–¶

With ~28.8 GB usable it runs up to ~43.5B at Q4 (~12.5B at FP16) at 8k context. Bigger spills into SSD swap and slows sharply.

How many tokens/sec does the Jetson Thor T3000 generate?â–¶

Generation is bandwidth-bound. At 273 GB/s an 8B model at Q4 runs ~44.5-53 tok/s (estimate).

Why does long context slow it down?â–¶

The KV cache for the context must also fit fast memory; once weights+KV exceed 28.8 GB it spills to SSD swap and throughput collapses. Each row shows the max usable context.