FU FuriosaAI Discrete GPU

FuriosaAI RNGD

Runs downloadable LLMs up to ~71B at Q4, ~217.6-264.5 tok/s on an 8B model.

FU
48GB · 1500 GB/s
48GB
VRAM
1500
GB/s bandwidth
71B
Max model (Q4)

What you can run

At 8k context, FP16 KV. Full-speed tok/s holds up to "max ctx"; beyond it the context spills to system RAM (CPU/PCIe offload) and drops to the degraded figure.

Model sizeFP16Q8Q4
3B 162.5-197.5 t/s
≤559k ctx · then ~4.3-5.3 (slow)
325-395 t/s
≤605k ctx · then ~8.7-10.5 (slow)
580.4-705.4 t/s
≤625k ctx · then ~15.5-18.8 (slow)
7-8B 60.9-74.1 t/s
≤203k ctx · then ~1.6-2 (slow)
121.9-148.1 t/s
≤264k ctx · then ~3.3-4 (slow)
217.6-264.5 t/s
≤291k ctx · then ~5.8-7.1 (slow)
13B 37.5-45.6 t/s
≤102k ctx · then ~1-1.2 (slow)
75-91.2 t/s
≤181k ctx · then ~2-2.4 (slow)
133.9-162.8 t/s
≤216k ctx · then ~3.6-4.3 (slow)
32B won't run 30.5-37 t/s
≤41k ctx · then ~0.8-1 (slow)
54.4-66.1 t/s
≤94k ctx · then ~1.5-1.8 (slow)
70B won't run won't run 24.9-30.2 t/s
≤11k ctx · then ~0.7-0.8 (slow)
120B won't run won't run won't run
400B won't run won't run won't run

Estimate: tok/s ≈ bandwidth ÷ (params × bytes/param) × efficiency (FuriosaAI). Cyan = fits ~44.2 GB usable. Beyond "ctx", KV spills to system RAM (CPU/PCIe offload) (rough degraded speed shown).

What else you need for a full build

Minimum supporting components for a desktop built around the FuriosaAI RNGD. Guidance — exact needs depend on your case, other parts and how much you offload to system RAM.

System RAM
≥ 64 GB · 128 GB recommended
At least the 48 GB of VRAM; 2× lets you CPU-offload models larger than VRAM.
Motherboard
1× free PCIe x16 slot (4.0/5.0)
A full-length x16 electrical slot; PCIe 4.0 is ample, 5.0 future-proofs.
CPU
Any modern multi-core
The GPU runs inference; more cores and memory channels only matter when offloading to RAM.
Storage
≥ 1 TB NVMe SSD
Weights are large — a 70B model at Q4 is ~40 GB; budget for a few.

Run popular models on the FuriosaAI RNGD

About the FuriosaAI RNGD

FuriosaAI RNGD is a discrete GPU from FuriosaAI. For local genAI the numbers that matter are its 48 GB (how big a model + context fits) and 1500 GB/s bandwidth (generation speed).

Frequently asked questions

What size LLM can the FuriosaAI RNGD run?

With ~44.2 GB usable it runs up to ~71B at Q4 (~20B at FP16) at 8k context. Bigger spills into system RAM (CPU/PCIe offload) and slows sharply.

How many tokens/sec does the FuriosaAI RNGD generate?

Generation is bandwidth-bound. At 1500 GB/s an 8B model at Q4 runs ~217.6-264.5 tok/s (estimate).

Why does long context slow it down?

The KV cache for the context must also fit fast memory; once weights+KV exceed 44.2 GB it spills to system RAM (CPU/PCIe offload) and throughput collapses. Each row shows the max usable context.

What else do I need to run the FuriosaAI RNGD?

Besides the card: at least 48GB of system RAM (2× the 48GB VRAM is ideal for CPU offload), a free PCIe x16 slot, any modern multi-core CPU, and a ≥1TB NVMe SSD for model weights.