Layered summaries keep context at 8k, yet Qwen degrades after ~120 chat turns
Zeeplankton · reddit · 2026-09-09
A developer on Reddit shares their conversation compaction setup: recursively generated hierarchical summaries (L1 factual extraction → L2 more narrative synthesis → L3), plus a tail of 10 raw timestamped messages, keeping context at around 5-8k tokens.
Yet after roughly 120-140 messages, Qwen Flash Next degrades severely—outputs quickly become incoherent and comically strange. The author notes 8k tokens should be trivial for Qwen, and the payload itself looks fairly coherent, so they suspect summaries themselves may pollute model coherence, and ask whether others are working on compaction.
More from coding & agent
- Nous Research pitches Hermes Agent as the only way to truly own your AI stack — Teknium · 2026-09-09
- One Person + Codex + Apify + Apollo: The Exact Stack for AI-Powered Outbound That Isn't Spam — aryanXmahajan · 2026-09-09
- GitHub Details 4 Changes That Cut Copilot's AI Cost Without Hurting Task Quality — marlene_zw · 2026-09-09
- Mac MCP 2.0 open-sources 81 tools giving AI agents local macOS control and background browser automation — bulutarkan · 2026-09-09
- Teknium: Ran Hermes 19 Hours with 1,390 Subagents and Zero Failures — Teknium · 2026-09-09
- Lessons from building a local macOS MCP server for ChatGPT-assisted coding and admin work — bulutarkan · 2026-09-09