128k context with compaction beats raw 256k/1M for quantized models

Informal-Trouble2183 · reddit · 2026-09-21

A Reddit user argues that quantization error accumulates across context length in quantized (around Q4) models, degrading quality on long contexts. They propose 128k as the sweet spot: pairing it with intelligent compaction (e.g. PI agent's compaction) preserves enough task-relevant information while avoiding quant-error accumulation, outperforming raw 256k or 1M contexts. Compaction also adds intelligent noise filtering — dropping file dumps, terminal logs, and outdated information. The post is a hypothesis seeking community experience rather than a systematic benchmark.

Original post →

More from coding & agent

coding & agent channel →