Paper: LLM swarms fail at 'recursive social improvement' — copying peers hurts exploration
maxhkw · x · 2026-10-03
A new paper by Kunal Jha, Max Kleiman-Weiner and Natasha Jaques (arXiv:2609.38516) tests whether self-improving LLM agents in swarms can learn from each other — "recursive social improvement." In controlled settings with a shared token budget, classic social-learning algorithms benefit from peers, but three LLMs don't: they earn less reward per token than solo learners and either explore too narrowly or run out of tokens before acting. Letting models write their own skills changes how they improve, but none beats independent learners at equal cost; peer copying concentrates the population around fewer discoveries, trading exploration for imitation.
More from coding & agent
- One image prompt gets Claude to build a playable browser game plus its trailer — justin_hart · 2026-10-03
- Dev shows ghost-zombie FPS alpha where all assets were built by an agent using Opus 5.5 + Blender — majidmanzarpour · 2026-10-03
- Ofir Press defends new bug-finding benchmark: training the behavior is fine, test-set contamination is not — OfirPress · 2026-10-03
- Dev argues AI is great at assets and code but bad at designing game mechanics that feel good — rms80 · 2026-10-03
- Free one-day curriculum takes you from AI agent basics to MCP and agent security — ifioknkem · 2026-10-03
- Using System One models in Swift: fast, deterministic decisions via Apple Foundation Models — rxwei · 2026-10-03