New: connect Claude & other AIs to GenAIList over MCP — research the catalog and contribute to the shared knowledge base. Learn how →
Huawei Ascend 950DT
Runs downloadable LLMs up to ~226B at Q4, ~580.4-705.4 tok/s on an 8B model.
What you can run
At 8k context, FP16 KV. Full-speed tok/s holds up to "max ctx"; beyond it the context spills to system RAM (CPU/PCIe offload) and drops to the degraded figure.
| Model size | FP16 | Q8 | Q4 |
|---|---|---|---|
| 3B |
433.3-526.7 t/s ≤1907k ctx · then ~4.3-5.3 (slow) |
866.7-1053.3 t/s ≤1953k ctx · then ~8.7-10.5 (slow) |
1547.6-1881 t/s ≤1973k ctx · then ~15.5-18.8 (slow) |
| 7-8B |
162.5-197.5 t/s ≤877k ctx · then ~1.6-2 (slow) |
325-395 t/s ≤938k ctx · then ~3.3-4 (slow) |
580.4-705.4 t/s ≤965k ctx · then ~5.8-7.1 (slow) |
| 13B |
100-121.5 t/s ≤641k ctx · then ~1-1.2 (slow) |
200-243.1 t/s ≤720k ctx · then ~2-2.4 (slow) |
357.1-434.1 t/s ≤755k ctx · then ~3.6-4.3 (slow) |
| 32B |
40.6-49.4 t/s ≤256k ctx · then ~0.4-0.5 (slow) |
81.3-98.8 t/s ≤378k ctx · then ~0.8-1 (slow) |
145.1-176.3 t/s ≤431k ctx · then ~1.5-1.8 (slow) |
| 70B | won't run |
37.1-45.1 t/s ≤186k ctx · then ~0.4-0.5 (slow) |
66.3-80.6 t/s ≤280k ctx · then ~0.7-0.8 (slow) |
| 120B | won't run |
21.7-26.3 t/s ≤34k ctx · then ~0.2-0.3 (slow) |
38.7-47 t/s ≤195k ctx · then ~0.4-0.5 (slow) |
| 400B | won't run | won't run | won't run |
Estimate: tok/s ≈ bandwidth ÷ (params × bytes/param) × efficiency (Huawei). Cyan = fits ~132.5 GB usable. Beyond "ctx", KV spills to system RAM (CPU/PCIe offload) (rough degraded speed shown).
What else you need for a full build▶
Minimum supporting components for a desktop built around the Huawei Ascend 950DT. Guidance — exact needs depend on your case, other parts and how much you offload to system RAM.
Run popular models on the Huawei Ascend 950DT
About the Huawei Ascend 950DT
Huawei Ascend 950DT is a discrete GPU from Huawei, 2026. For local genAI the numbers that matter are its 144 GB (how big a model + context fits) and 4000 GB/s bandwidth (generation speed).
Frequently asked questions
What size LLM can the Huawei Ascend 950DT run?▶
With ~132.5 GB usable it runs up to ~226B at Q4 (~64B at FP16) at 8k context. Bigger spills into system RAM (CPU/PCIe offload) and slows sharply.
How many tokens/sec does the Huawei Ascend 950DT generate?▶
Generation is bandwidth-bound. At 4000 GB/s an 8B model at Q4 runs ~580.4-705.4 tok/s (estimate).
Why does long context slow it down?▶
The KV cache for the context must also fit fast memory; once weights+KV exceed 132.5 GB it spills to system RAM (CPU/PCIe offload) and throughput collapses. Each row shows the max usable context.
What else do I need to run the Huawei Ascend 950DT?▶
Besides the card: at least 144GB of system RAM (2× the 144GB VRAM is ideal for CPU offload), a free PCIe x16 slot, any modern multi-core CPU, and a ≥1TB NVMe SSD for model weights.