Strata engine runs Qwen3.8 Flash Next on a 12GB laptop at 51 t/s, 1500 t/s prefill

MLDataScientist · reddit · 2026-09-30

A hands-on test of the new Strata inference engine on a 12GB VRAM laptop (5070ti, 64GB RAM) running ISTA-DASLab's Qwen3.8 Flash Next GGUF: 51 t/s generation at 43k context (vs 23 t/s on stock llama.cpp with the same quant) and 1500 t/s prefill for 32k context (vs 100 t/s), with 11GB VRAM used. Caveats: Strata currently supports only this model and select ISTA-DASLab quants (IQ3S claimed to recover full coding performance), NVIDIA-only with experimental AMD support. Early kv-cache bugs are fixed.

Original post →

More from Infra

Infra channel →