Strata engine runs Qwen3.8 Flash Next on a 12GB laptop at 51 t/s, 1500 t/s prefill
MLDataScientist · reddit · 2026-09-30
A hands-on test of the new Strata inference engine on a 12GB VRAM laptop (5070ti, 64GB RAM) running ISTA-DASLab's Qwen3.8 Flash Next GGUF: 51 t/s generation at 43k context (vs 23 t/s on stock llama.cpp with the same quant) and 1500 t/s prefill for 32k context (vs 100 t/s), with 11GB VRAM used. Caveats: Strata currently supports only this model and select ISTA-DASLab quants (IQ3S claimed to recover full coding performance), NVIDIA-only with experimental AMD support. Early kv-cache bugs are fixed.
More from Infra
- 8% of Asia-to-US air freight is now data center parts — 30 full freighters a day — yacineMTB · 2026-09-30
- mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB — Spiritual_Impress_30 · 2026-09-30
- Hugging Face ships tokenizers v1, often tens of times faster than v0.23 — ariG23498 · 2026-09-30
- Altman says OpenAI's custom chip program comes online in H1 2027, betting on inference advantage — johncoogan · 2026-09-30
- Rumor: DeepSeek's rumored single-GPU model may have been trained on Ascend — teortaxesTex · 2026-09-30
- AAOI burnt capex on US vertical integration, CW laser yield still poor — jwt0625 · 2026-09-30