UT Austin study: context compression tuned to cut tokens can make coding agents 20-80% slower
omarsar0 · x · 2026-10-05
A UT Austin study on context compression in coding agents ran nearly 35,000 agent runs on SWE-bench Verified and Terminal-Bench, independently varying how context is compressed, when compression triggers, and how much is removed.
Key findings:
- Token savings aren't time savings: on Terminal-Bench with Qwen, policies using about a third of the tokens take 20% to 80% longer than keeping full context
- Trigger trade-offs: step-triggered policies cut the most tokens per step but need 10-27% more model calls; threshold-triggered ones cut 22-55% with call counts close to full context
- Policies don't transfer across models: what works for Qwen drops Devstral to 38.7% and makes it slower
If your compaction policy is tuned only to cut tokens, it may be slowing your agent down — measure latency.
More from coding & agent
- Ponytail Hits 150k+ Stars Telling AI Coding Agents to Stop Overbuilding — we93 · 2026-10-05
- Dev predicts companies will wall off MCP servers, ushering in agent-to-agent era — JosephJacks_ · 2026-10-05
- Dots and Muse AI Coding Tools Can't Resize Browser Below ~500px, Blocking Mobile Testing — pkragthorpe · 2026-10-05
- Dev orchestrates 6 AI skills for a test-fix pipeline, spending just $3/week on DeepSeek — Own-Awareness8037 · 2026-10-05
- fastapi-gql-mcp: one GraphQL MCP surface cuts tool tokens from ~11k to ~2.5k and stops payload bloat — tangkikodo · 2026-10-05
- Tencent Hunyuan's RSR Boosts 27B Model Terminal-Bench 2 pass@3 from 57% to 74% — Tencent-Hunyuan · 2026-10-05