Debugging OpenCode stalls: Qwen3.6 hybrid memory forces llama-server full prompt re-processing

MysteriousInterest32 · reddit · 2026-10-05

Running Qwen3.6 35B-A3B (Q4KM) via llama-server on a 16GB 4060 Ti with most experts on CPU, the author found every OpenCode turn began with a long pause that grew with the session.

A valuable gotcha log for anyone running hybrid-architecture local models in coding-agent workflows.

Original post →

More from coding & agent

coding & agent channel →