HU Huawei Discrete GPU

Huawei Ascend 950DT

Runs downloadable LLMs up to ~226B at Q4, ~580.4-705.4 tok/s on an 8B model.

HU
144GB · 4000 GB/s
144GB
VRAM
4000
GB/s bandwidth
226B
Max model (Q4)

What you can run

At 8k context, FP16 KV. Full-speed tok/s holds up to "max ctx"; beyond it the context spills to system RAM (CPU/PCIe offload) and drops to the degraded figure.

Model sizeFP16Q8Q4
3B 433.3-526.7 t/s
≤1907k ctx · then ~4.3-5.3 (slow)
866.7-1053.3 t/s
≤1953k ctx · then ~8.7-10.5 (slow)
1547.6-1881 t/s
≤1973k ctx · then ~15.5-18.8 (slow)
7-8B 162.5-197.5 t/s
≤877k ctx · then ~1.6-2 (slow)
325-395 t/s
≤938k ctx · then ~3.3-4 (slow)
580.4-705.4 t/s
≤965k ctx · then ~5.8-7.1 (slow)
13B 100-121.5 t/s
≤641k ctx · then ~1-1.2 (slow)
200-243.1 t/s
≤720k ctx · then ~2-2.4 (slow)
357.1-434.1 t/s
≤755k ctx · then ~3.6-4.3 (slow)
32B 40.6-49.4 t/s
≤256k ctx · then ~0.4-0.5 (slow)
81.3-98.8 t/s
≤378k ctx · then ~0.8-1 (slow)
145.1-176.3 t/s
≤431k ctx · then ~1.5-1.8 (slow)
70B won't run 37.1-45.1 t/s
≤186k ctx · then ~0.4-0.5 (slow)
66.3-80.6 t/s
≤280k ctx · then ~0.7-0.8 (slow)
120B won't run 21.7-26.3 t/s
≤34k ctx · then ~0.2-0.3 (slow)
38.7-47 t/s
≤195k ctx · then ~0.4-0.5 (slow)
400B won't run won't run won't run

Estimate: tok/s ≈ bandwidth ÷ (params × bytes/param) × efficiency (Huawei). Cyan = fits ~132.5 GB usable. Beyond "ctx", KV spills to system RAM (CPU/PCIe offload) (rough degraded speed shown).

What else you need for a full build

Minimum supporting components for a desktop built around the Huawei Ascend 950DT. Guidance — exact needs depend on your case, other parts and how much you offload to system RAM.

System RAM
≥ 256 GB · 512 GB recommended
At least the 144 GB of VRAM; 2× lets you CPU-offload models larger than VRAM.
Motherboard
1× free PCIe x16 slot (4.0/5.0)
A full-length x16 electrical slot; PCIe 4.0 is ample, 5.0 future-proofs.
CPU
Any modern multi-core
The GPU runs inference; more cores and memory channels only matter when offloading to RAM.
Storage
≥ 1 TB NVMe SSD
Weights are large — a 70B model at Q4 is ~40 GB; budget for a few.

Run popular models on the Huawei Ascend 950DT

About the Huawei Ascend 950DT

Huawei Ascend 950DT is a discrete GPU from Huawei, 2026. For local genAI the numbers that matter are its 144 GB (how big a model + context fits) and 4000 GB/s bandwidth (generation speed).

Frequently asked questions

What size LLM can the Huawei Ascend 950DT run?

With ~132.5 GB usable it runs up to ~226B at Q4 (~64B at FP16) at 8k context. Bigger spills into system RAM (CPU/PCIe offload) and slows sharply.

How many tokens/sec does the Huawei Ascend 950DT generate?

Generation is bandwidth-bound. At 4000 GB/s an 8B model at Q4 runs ~580.4-705.4 tok/s (estimate).

Why does long context slow it down?

The KV cache for the context must also fit fast memory; once weights+KV exceed 132.5 GB it spills to system RAM (CPU/PCIe offload) and throughput collapses. Each row shows the max usable context.

What else do I need to run the Huawei Ascend 950DT?

Besides the card: at least 144GB of system RAM (2× the 144GB VRAM is ideal for CPU offload), a free PCIe x16 slot, any modern multi-core CPU, and a ≥1TB NVMe SSD for model weights.