New: connect Claude & other AIs to GenAIList over MCP — research the catalog and contribute to the shared knowledge base. Learn how →
FuriosaAI RNGD
Runs downloadable LLMs up to ~71B at Q4, ~217.6-264.5 tok/s on an 8B model.
What you can run
At 8k context, FP16 KV. Full-speed tok/s holds up to "max ctx"; beyond it the context spills to system RAM (CPU/PCIe offload) and drops to the degraded figure.
| Model size | FP16 | Q8 | Q4 |
|---|---|---|---|
| 3B |
162.5-197.5 t/s ≤559k ctx · then ~4.3-5.3 (slow) |
325-395 t/s ≤605k ctx · then ~8.7-10.5 (slow) |
580.4-705.4 t/s ≤625k ctx · then ~15.5-18.8 (slow) |
| 7-8B |
60.9-74.1 t/s ≤203k ctx · then ~1.6-2 (slow) |
121.9-148.1 t/s ≤264k ctx · then ~3.3-4 (slow) |
217.6-264.5 t/s ≤291k ctx · then ~5.8-7.1 (slow) |
| 13B |
37.5-45.6 t/s ≤102k ctx · then ~1-1.2 (slow) |
75-91.2 t/s ≤181k ctx · then ~2-2.4 (slow) |
133.9-162.8 t/s ≤216k ctx · then ~3.6-4.3 (slow) |
| 32B | won't run |
30.5-37 t/s ≤41k ctx · then ~0.8-1 (slow) |
54.4-66.1 t/s ≤94k ctx · then ~1.5-1.8 (slow) |
| 70B | won't run | won't run |
24.9-30.2 t/s ≤11k ctx · then ~0.7-0.8 (slow) |
| 120B | won't run | won't run | won't run |
| 400B | won't run | won't run | won't run |
Estimate: tok/s ≈ bandwidth ÷ (params × bytes/param) × efficiency (FuriosaAI). Cyan = fits ~44.2 GB usable. Beyond "ctx", KV spills to system RAM (CPU/PCIe offload) (rough degraded speed shown).
What else you need for a full build▶
Minimum supporting components for a desktop built around the FuriosaAI RNGD. Guidance — exact needs depend on your case, other parts and how much you offload to system RAM.
Run popular models on the FuriosaAI RNGD
About the FuriosaAI RNGD
FuriosaAI RNGD is a discrete GPU from FuriosaAI. For local genAI the numbers that matter are its 48 GB (how big a model + context fits) and 1500 GB/s bandwidth (generation speed).
Frequently asked questions
What size LLM can the FuriosaAI RNGD run?▶
With ~44.2 GB usable it runs up to ~71B at Q4 (~20B at FP16) at 8k context. Bigger spills into system RAM (CPU/PCIe offload) and slows sharply.
How many tokens/sec does the FuriosaAI RNGD generate?▶
Generation is bandwidth-bound. At 1500 GB/s an 8B model at Q4 runs ~217.6-264.5 tok/s (estimate).
Why does long context slow it down?▶
The KV cache for the context must also fit fast memory; once weights+KV exceed 44.2 GB it spills to system RAM (CPU/PCIe offload) and throughput collapses. Each row shows the max usable context.
What else do I need to run the FuriosaAI RNGD?▶
Besides the card: at least 48GB of system RAM (2× the 48GB VRAM is ideal for CPU offload), a free PCIe x16 slot, any modern multi-core CPU, and a ≥1TB NVMe SSD for model weights.