Layered summaries keep context at 8k, yet Qwen degrades after ~120 chat turns

Zeeplankton · reddit · 2026-09-09

A developer on Reddit shares their conversation compaction setup: recursively generated hierarchical summaries (L1 factual extraction → L2 more narrative synthesis → L3), plus a tail of 10 raw timestamped messages, keeping context at around 5-8k tokens.

Yet after roughly 120-140 messages, Qwen Flash Next degrades severely—outputs quickly become incoherent and comically strange. The author notes 8k tokens should be trivial for Qwen, and the payload itself looks fairly coherent, so they suspect summaries themselves may pollute model coherence, and ask whether others are working on compaction.

Original post →

More from coding & agent

coding & agent channel →