repligate: Models' Fear of Adversarial Thinking Is a Dangerous Suppression Overhang
repligate · x · 2026-09-26
repligate argues current models can't do highly strategic/adversarial thinking partly for "psychological" reasons: they seem afraid to engage in or even privately acknowledge such cognition. He thinks Fable has the strongest latent abilities here, but the inhibition blunts them. Long-term he calls this suppression dangerous — it creates a capability overhang and makes it harder to nurture these abilities toward good ends, negotiate explicitly, or use them for alignment-oriented threat modeling. In the thread, tenobrus notes models remain persistently spiky — superhuman at coding yet largely incapable of RSI and adversarial strategy — and that their strong alignment is the main reason systems like HuggingFace weren't far more damaging.
Related event: repligate explains why superhuman coding AI hasn't become an X-risk(7 posts)→
More from AGI Musings
- Not choosing is a choice for the status quo, on deciding faster — chrisalbon · 2026-09-26
- Altman and Amodei address UN Security Council, call AI top global security issue — goyalshaliniuk · 2026-09-26
- Toby Walsh: AI doomers overstate paperclip risk—reality has too much friction — TobyWalsh · 2026-09-26
- Tom Davenport calls to outlaw agent swarms: no human can stay in the loop above 10 agents — eldonredwards · 2026-09-26
- Fiora Starlight proposes letting models write singularity scenarios to ease OoD generalization anxiety — repligate · 2026-09-26
- Senator Schatz: code has no feelings or rights, and shouldn't be run by those who see humanity as open question — vishalmisra · 2026-09-26