Context Language Models: Letting LLMs Rewrite Their Own Context Cuts FLOPs 21.5% While Boosting Accuracy 11.4%
RulinShao · x · 2026-10-05
Researchers from UW, Meta Superintelligence Labs, MIT and others introduce Context Language Models (CLMs): instead of relying on an external harness, the model treats its context as a file and directly rewrites it, learning when to compress, delete, preserve or reorganize information — naturally extending to multi-agent setups where agent contexts coexist as files.
Key results:
- Zero-shot, CLMs beat SOTA context-management strategies: +11.4% accuracy with 21.5% fewer FLOPs on BrowseComp-Plus; +7% scores with 59% fewer FLOPs on 12-hour EdgeBench; 65% greater improvement at matched compute on a 24-hour multi-repo agent-swarm task.
- Steering CLMs with natural-language instructions evolved via a standard skill-optimization loop improves held-out accuracy by up to 35.9 points.
- An online RL method lifts Qwen3.5-9B by 47.6% on BrowseComp-Plus with 12% fewer FLOPs.
- A co-designed Suffix Cache Reuse serving scheme cuts server-side compute by 35% vs. standard SGLang at matched performance.
Authors include Luke Zettlemoyer, Pang Wei Koh, and Nathan Lambert.
More from coding & agent
- Apple quietly shipped MagSafe 3 cable firmware 3.2.0; GLM used to diff the changes — steipete · 2026-10-05
- Databricks-backed Omnigent hits 10.5k GitHub stars: open-source meta-harness to swap agent harnesses without rewrites — bibryam · 2026-10-05
- He rebuilt his personal site with an AI agent in 5 minutes, migrating 83 essays and saving $120/year — thisiskp_ · 2026-10-05
- Agent policy-boundary starter updated with idempotency, audit logging and independent validation — Dapper-Roof2370 · 2026-10-05
- Biggest multi-agent system win was making the agents replaceable, not central — Druss_ · 2026-10-05
- Anthropic exec hits inbox zero for the first time by making it an explicit goal for his Claude agent — BorisMPower · 2026-10-05