GM GMKtec Unified memory

GMKtec EVO-X3

Runs downloadable LLMs up to ~195B at Q4, ~37.1-45.1 tok/s on an 8B model.

GM
128GB · 256 GB/s
128GB
Unified memory
256
GB/s bandwidth
195B
Max model (Q4)

What you can run

At 8k context, FP16 KV. Full-speed tok/s holds up to "max ctx"; beyond it the context spills to SSD swap and drops to the degraded figure.

Model sizeFP16Q8Q4
3B 27.7-33.7 t/s
≤1643k ctx · then ~1.1-1.3 (slow)
55.5-67.4 t/s
≤1689k ctx · then ~2.2-2.6 (slow)
99-120.4 t/s
≤1709k ctx · then ~3.9-4.7 (slow)
7-8B 10.4-12.6 t/s
≤745k ctx · then ~0.4-0.5 (slow)
20.8-25.3 t/s
≤806k ctx · then ~0.8-1 (slow)
37.1-45.1 t/s
≤833k ctx · then ~1.5-1.8 (slow)
13B 6.4-7.8 t/s
≤535k ctx · then ~0.3-0.3 (slow)
12.8-15.6 t/s
≤615k ctx · then ~0.5-0.6 (slow)
22.9-27.8 t/s
≤650k ctx · then ~0.9-1.1 (slow)
32B 2.6-3.2 t/s
≤190k ctx · then ~0.1-0.1 (slow)
5.2-6.3 t/s
≤312k ctx · then ~0.2-0.2 (slow)
9.3-11.3 t/s
≤365k ctx · then ~0.4-0.4 (slow)
70B won't run 2.4-2.9 t/s
≤133k ctx · then ~0.1-0.1 (slow)
4.2-5.2 t/s
≤227k ctx · then ~0.2-0.2 (slow)
120B won't run won't run 2.5-3 t/s
≤142k ctx · then ~0.1-0.1 (slow)
400B won't run won't run won't run

Estimate: tok/s ≈ bandwidth ÷ (params × bytes/param) × efficiency (GMKtec). Cyan = fits ~115.2 GB usable. Beyond "ctx", KV spills to SSD swap (rough degraded speed shown).

Run popular models on the GMKtec EVO-X3

About the GMKtec EVO-X3

GMKtec EVO-X3 is a unified-memory system from GMKtec, 2026. For local genAI the numbers that matter are its 128 GB (how big a model + context fits) and 256 GB/s bandwidth (generation speed).

Frequently asked questions

What size LLM can the GMKtec EVO-X3 run?â–¶

With ~115.2 GB usable it runs up to ~195B at Q4 (~55.5B at FP16) at 8k context. Bigger spills into SSD swap and slows sharply.

How many tokens/sec does the GMKtec EVO-X3 generate?â–¶

Generation is bandwidth-bound. At 256 GB/s an 8B model at Q4 runs ~37.1-45.1 tok/s (estimate).

Why does long context slow it down?â–¶

The KV cache for the context must also fit fast memory; once weights+KV exceed 115.2 GB it spills to SSD swap and throughput collapses. Each row shows the max usable context.