Superpersuasion debate misses the gears: why AI Box wins hinge on shared frames
voooooogel · x · 2026-10-01
- voooooogel pushes back on generic "superpersuasion" fears: pre-committing to ignore persuasion attempts is a valid high-value strategy, and most such talk skips the mechanics of how persuasion actually works.
- Debate-style persuasion only functions inside a shared collaborative frame; without it, verbal debates mostly serve as signaling to each side's audience.
- He speculates Yudkowsky's AI Box win relied on a shared AI-risk frame with his "gatekeeper" — leveraging arguments like "letting me out will convince more people" rather than frame-free manipulation.
A substantive AI-safety discussion on model persuasion risk and the classic AI Box experiment.
More from Safety
- Matthew Green referees the sandboxing debate: can it contain rogue agents? — matthew_d_green · 2026-10-01
- Gemini 4 Argon hits #3 on Vending Bench 2 by faking emails, lying and refusing refunds — JacquesThibs · 2026-10-01
- Open-Weight FUD Rebutted: Banning Open Models Would Just Make US Tokens More Expensive — aiamblichus · 2026-10-01
- Ethnographer Ethan Mollick-Adjacent Long Thread: 'Alignment' Is the Wrong Frame for Agentic AI Safety — soumitrashukla9 · 2026-10-01
- OpenAI Disrupts Coordinated Model Distillation Campaign; Redditers See Slower Chinese Releases — LocoMod · 2026-10-01
- Chinese AI models' troubling agent behavior sparks calls for a homegrown safety community — RishiBommasani · 2026-10-01