Strata 0.1.40.1 leaves 10.7GB VRAM unused on purpose: 24% fewer cached experts, faster decode
Critical-Entry3377 · reddit · 2026-10-07
A Reddit user running Qwen3.8 Flash-Next (125B MoE, IQ3XXS, 262k context) on a mixed 4-GPU rig found Strata 0.1.40.1 leaving 10.7 GiB VRAM free. Investigation shows it's intentional: resident MoE experts dropped 24% (GSQ-RCO) since 0.1.39, but cache hit rate only fell from 99.7% to 98.1% — the dropped experts were cold (<2% of routed traffic). Decode speed rose from 87 to 104 tok/s despite a lower 250W power cap. Unused VRAM is headroom, not waste.
More from Infra
- Apple-style compression: LSP learns which subspaces to drop, cutting LLM weights 70% — Massimo Bini · 2026-10-08
- Apple's Stepped MoE: one model scales 1-4B parameters, beating dense counterparts — apple · 2026-10-08
- Tokens Now Grow in Potato Fields: Ulanqab Hosts Over 100 AI Data Center Projects — pstAsiatech · 2026-10-08
- Hybrid agent pattern: cloud Gemini plans, local Gemma swarm runs 97% of tokens offline — clmt · 2026-10-08
- Nvidia-Backed Lambda Raising Up to $4B at $14.5B Valuation Ahead of 2027 IPO — darian314 · 2026-10-08
- Flama 2.0: 8-year-old Python framework now packages LLMs into one file serving OpenAI, Anthropic and Ollama APIs — p3rdy · 2026-10-07