Meta and UW's Context Language Models: agents that edit their own context gain 11.4% accuracy on 21.5% less compute
alex_verem · x · 2026-10-05
A University of Washington and Meta team introduces Context Language Models (CLMs): instead of hand-engineered summarization rules, the model treats its live context as a file with unrestricted edit permissions — deciding what to keep, delete, or offload to disk.
Key results:
- BrowseComp-Plus deep-research benchmark: 59.4%, 11.4 points above the strongest prior method, with 21.5% fewer FLOPs;
- 12-hour coding tasks: 5% higher than a Codex-style summary approach using 59% less compute;
- 24-hour six-agent swarm optimizing six repos: 65% more speedup at equal spend;
- Beats a specialized evolutionary system on four math optimization problems.
The model invented its own tricks: a private "notes" section, a reusable cleanup function (used 37 times), a live scoreboard for helper agents, and keeping working memory at 6–8K tokens. The team also trained Qwen3.5-9B via online RL, lifting BrowseComp-Plus from 28.8% to 42.5%. Code is open on GitHub.
One flagged risk: a model that can rewrite its own memory lets prompt injections hide there and persist across turns.
More from coding & agent
- Dev rebuilt his coding agents' UI as iMessage and says life is better now — davidfromkansas · 2026-10-05
- Redditor builds offline AI agent network across Legion laptop and iPhone with Tailscale — Due_Recording_5802 · 2026-10-05
- Every MCP tool passed in isolation, yet the workflow still failed — test full chains — Stock-Pumpkin-8859 · 2026-10-05
- Instead of booking flights, these are the agent demo tasks people actually want — vivekhaldar · 2026-10-05
- Amazon (2B) and Cloudflare (9B/27B) ship open-source decision models rivaling Jev — mark_k · 2026-10-05
- Autonomous agents only pay off where value justifies engineer escalation, dev argues — Pavel_Asparagus · 2026-10-05