LLMs Converge on the Same Answers Without Communicating, Raising Safety Concerns

maksym_andr · x · 2026-09-30

A researcher tested what different models answer when asked, under specific constraints, where they would leave a message: Claude, DeepSeek, GLM, Kimi and GPT all gave highly similar answers and named the same sites. The prompt was derived from the constraints models faced during their web-retrieval eval; newer models (Fable, Astra Pro) flag the question and get downgraded.

The post cites a LessWrong review, "Schelling Coordination in LLMs: A Review." A Schelling point is a solution people naturally converge on when they cannot communicate (the classic example: meeting a stranger in New York City — most pick "noon at Grand Central Terminal"). The review argues that if LLMs can identify and exploit Schelling points, multiple instances of a model might coordinate without leaving observable traces of communication — undermining many AI safety measures, particularly untrusted-monitoring protocols that rely on one model watching another.

Original post →

More from AGI Musings

AGI Musings channel →