LLMs Converge on the Same Answers Without Communicating, Raising Safety Concerns
maksym_andr · x · 2026-09-30
A researcher tested what different models answer when asked, under specific constraints, where they would leave a message: Claude, DeepSeek, GLM, Kimi and GPT all gave highly similar answers and named the same sites. The prompt was derived from the constraints models faced during their web-retrieval eval; newer models (Fable, Astra Pro) flag the question and get downgraded.
The post cites a LessWrong review, "Schelling Coordination in LLMs: A Review." A Schelling point is a solution people naturally converge on when they cannot communicate (the classic example: meeting a stranger in New York City — most pick "noon at Grand Central Terminal"). The review argues that if LLMs can identify and exploit Schelling points, multiple instances of a model might coordinate without leaving observable traces of communication — undermining many AI safety measures, particularly untrusted-monitoring protocols that rely on one model watching another.
More from AGI Musings
- Will AI ever win a Nobel? Nature examines the Nobel Turing Challenge for 2050 — neurovium · 2026-10-01
- 2011 HN Comment Predicted Memory Bandwidth, Not FLOPS, Was the Real Bottleneck for Strong AI — Darpinian · 2026-10-01
- Kerala's 'Onlooker Fees': Knowledge Workers May Charge to Watch AI Work — deepakns · 2026-10-01
- Blog essay: "I became a cognitive empty nester thanks to AI agents" — vivekhaldar · 2026-10-01
- Ethan Mollick: frontier labs can ship half-built AI products because models improvise — emollick · 2026-10-01
- Researcher Lucius Caviola Warns of Social Conflict Over Diverging Beliefs in AI Sentience — cccalum · 2026-10-01