Debugging OpenCode stalls: Qwen3.6 hybrid memory forces llama-server full prompt re-processing
MysteriousInterest32 · reddit · 2026-10-05
Running Qwen3.6 35B-A3B (Q4KM) via llama-server on a 16GB 4060 Ti with most experts on CPU, the author found every OpenCode turn began with a long pause that grew with the session.
- Initial suspects — PCIe 4.0 x8 bandwidth and CPU↔GPU expert copies — only partially explained it; raising -ub helped slightly, and a bigger card shortened but didn't eliminate the stalls.
- Server logs revealed most turns forced full prompt re-processing: Qwen3.6's hybrid/recurrent memory can't be rolled back partway, so any request mismatch without an earlier checkpoint means restarting from token one.
- The chat template was dropping earlier turns' thinking, changing every request. Passing preservethinking through template kwargs stopped most re-processing.
- Trade-off: keeping thinking fills context faster, so compaction runs more often — and compaction summaries never hit the cache, still causing full re-processing.
A valuable gotcha log for anyone running hybrid-architecture local models in coding-agent workflows.
More from coding & agent
- The viral vibe-coder SEO prompt: AI fixes the tech, but backlinks still take grinding — sujingshen · 2026-10-05
- Blender + Claude + Polyxd: swapping prompts for sliders to drive a 3D island scene live — sidahuj · 2026-10-05
- Borrowing aviation's controlled English (ASD-STE100) to strip AI fluff and fake completions — sujingshen · 2026-10-05
- Use Magpie CLI to check quotas and route sub-agents by urgency to save tokens — lxfater · 2026-10-05
- Give Claude real designer assets: 40 vetted resources to fix vibe-coded UI, says founder — PrajwalTomar_ · 2026-10-05
- Claude Opus 5.5 builds an interactive lens lab in one 1h26m shot for $25.66 — nikola_mr64990 · 2026-10-05