// RUN โ€” FITS

Can the NVIDIA A100 SXM4 80GB run TurboVLA?

Yes โ€” here's how fast and up to what context, with real measured numbers.

TurboVLA on the NVIDIA A100 SXM4 80GB

QuantFits?Max contextSpeedBeyond context
FP16 yes 1094k 3721.2-4434.8 t/s ~73-87 (slow)
Q8 yes 1097k 7442.4-8869.7 t/s ~146-174 (slow)
Q4 yes 1098k 13289.9-15838.7 t/s ~260.7-310.7 (slow)

Estimate (memory-bound). Beyond "max context" the KV cache spills to system RAM offload (depends on your system RAM) and speed drops to the "beyond" figure.

FAQ

Can the NVIDIA A100 SXM4 80GB run TurboVLA?โ–ถ

Yes. At Q4 it generates ~13289.9-15838.7 tok/s and fits up to 1098k context; beyond that it spills to system RAM offload (depends on your system RAM) and slows to ~260.7-310.7 tok/s.

How much context fits?โ–ถ

At Q4, up to about 1098k tokens stay in fast memory; longer context spills to system RAM offload (depends on your system RAM) and throughput collapses.