Why specialized agents that update their own weights beat generalist models

willcb · x · 2026-09-26

The author argues that expecting models to be superhuman at everything lacks historical or theoretical precedent — the smartest humans are exceptionally specialized. A long-running adaptive agent harness is a primitive "specialized agent" limited by context length, filesystem expressivity, and bespoke retrieval; letting the agent change its own weights is strictly more powerful. The bitter-lesson version: use RL to train models to update themselves on-the-fly however they best see fit, just as compaction and subagent delegation are trained end-to-end today.

Related event: Researchers Debate Whether Task-Specific RL Can Produce General Agents(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →