Qwen 3.8 Flash Next q2_0 hits 10 tok/s on a 6GB VRAM laptop via Strata engine
dampflokfreund · reddit · 2026-10-03
- Redditor dampflokfreund ran Qwen 3.8 Flash Next q20 (120B params) on an RTX 2060 laptop (32GB RAM + 6GB VRAM) using the Strata engine
- Result: 10 tok/s decode at 50K context and 100 tok/s prefill — far beyond expectations for such hardware
- Command used: ./START-HERE.bat --draft-vocab en --vram-reserve-mib 100 --kv q40
- Tradeoffs: 4-bit KV cache, no vision; prefill latency remains the bottleneck for long prompts
- Compared to Qwen A35B A3B (5x faster prefill, 100K context without KV quant), it's not a full replacement yet, but an impressive showing
More from Infra
- Chip design is a ~10^2,632,341 search problem — AI and agents are turning hardware into search — ai · 2026-10-03
- Redis Creator antirez Ships DwarfStar 4, an MIT-Licensed Local Inference Engine — petrusenko_max · 2026-10-03
- The $899 Mac Mini M6 was built for 24/7 agentic computing — most owners use it to check email — hey_abusiddik · 2026-10-03
- Modal VM Sandboxes hit GA: demo runs Docker Compose apps, tests and coding agents — charles_irl · 2026-10-03
- Community poll of 793 MLX users: oMLX wins at 55.5%, dominating Ultra chips — HankYeomans · 2026-10-03
- Local AI user asks: why does everyone enjoy taming the beast of complexity? — kathi7 · 2026-10-03