Paper: LLM swarms fail at 'recursive social improvement' — copying peers hurts exploration

maxhkw · x · 2026-10-03

A new paper by Kunal Jha, Max Kleiman-Weiner and Natasha Jaques (arXiv:2609.38516) tests whether self-improving LLM agents in swarms can learn from each other — "recursive social improvement." In controlled settings with a shared token budget, classic social-learning algorithms benefit from peers, but three LLMs don't: they earn less reward per token than solo learners and either explore too narrowly or run out of tokens before acting. Letting models write their own skills changes how they improve, but none beats independent learners at equal cost; peer copying concentrates the population around fewer discoveries, trading exploration for imitation.

Original post →

More from coding & agent

coding & agent channel →