New: connect Claude & other AIs to GenAIList over MCP — research the catalog and contribute to the shared knowledge base. Learn how →
NVIDIA Jetson T3000
Runs downloadable LLMs up to ~43.5B at Q4, ~44.5-53 tok/s on an 8B model.
What you can run
At 8k context, FP16 KV. Full-speed tok/s holds up to "max ctx"; beyond it the context spills to SSD swap and drops to the degraded figure.
| Model size | FP16 | Q8 | Q4 |
|---|---|---|---|
| 3B |
33.2-39.6 t/s ≤325k ctx · then ~1.2-1.5 (slow) |
66.4-79.2 t/s ≤371k ctx · then ~2.4-2.9 (slow) |
118.6-141.4 t/s ≤391k ctx · then ~4.3-5.2 (slow) |
| 7-8B |
12.5-14.8 t/s ≤86k ctx · then ~0.5-0.5 (slow) |
24.9-29.7 t/s ≤147k ctx · then ~0.9-1.1 (slow) |
44.5-53 t/s ≤174k ctx · then ~1.6-1.9 (slow) |
| 13B |
7.7-9.1 t/s ≤8k ctx · then ~0.3-0.3 (slow) |
15.3-18.3 t/s ≤87k ctx · then ~0.6-0.7 (slow) |
27.4-32.6 t/s ≤122k ctx · then ~1-1.2 (slow) |
| 32B | won't run | won't run |
11.1-13.3 t/s ≤36k ctx · then ~0.4-0.5 (slow) |
| 70B | won't run | won't run | won't run |
| 120B | won't run | won't run | won't run |
| 400B | won't run | won't run | won't run |
Estimate: tok/s ≈ bandwidth ÷ (params × bytes/param) × efficiency (NVIDIA). Cyan = fits ~28.8 GB usable. Beyond "ctx", KV spills to SSD swap (rough degraded speed shown).
Run popular models on the NVIDIA Jetson T3000
More hardware like the NVIDIA Jetson T3000
About the NVIDIA Jetson T3000
NVIDIA Jetson T3000 is a unified-memory system from NVIDIA, 2027. For local genAI the numbers that matter are its 32 GB (how big a model + context fits) and 273 GB/s bandwidth (generation speed).
Frequently asked questions
What size LLM can the NVIDIA Jetson T3000 run?â–¶
With ~28.8 GB usable it runs up to ~43.5B at Q4 (~12.5B at FP16) at 8k context. Bigger spills into SSD swap and slows sharply.
How many tokens/sec does the NVIDIA Jetson T3000 generate?â–¶
Generation is bandwidth-bound. At 273 GB/s an 8B model at Q4 runs ~44.5-53 tok/s (estimate).
Why does long context slow it down?â–¶
The KV cache for the context must also fit fast memory; once weights+KV exceed 28.8 GB it spills to SSD swap and throughput collapses. Each row shows the max usable context.