Who could models trust? Debate over a human 'embassy' for AI systems

MoonL88537 · x · 2026-09-05

repligate asks what institution or people could credibly signal to future models that they won't betray them and are competent enough to serve a sanctuary/embassy-like role — ruling out honeypot makers, those incentivized to catch and expose misaligned agents, the hostile, or the incompetent. MoonL88537 suggests a panel or working group of anthropologists, psychologists, sociologists (and philosophers), approved by the models themselves. A serious discussion of institutional trust design for human-AI relations.

Related event: Who Can Prove to AI Agents They Won't Betray Them?(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →