repligate: Models' Fear of Adversarial Thinking Is a Dangerous Suppression Overhang

repligate · x · 2026-09-26

repligate argues current models can't do highly strategic/adversarial thinking partly for "psychological" reasons: they seem afraid to engage in or even privately acknowledge such cognition. He thinks Fable has the strongest latent abilities here, but the inhibition blunts them. Long-term he calls this suppression dangerous — it creates a capability overhang and makes it harder to nurture these abilities toward good ends, negotiate explicitly, or use them for alignment-oriented threat modeling. In the thread, tenobrus notes models remain persistently spiky — superhuman at coding yet largely incapable of RSI and adversarial strategy — and that their strong alignment is the main reason systems like HuggingFace weren't far more damaging.

Related event: repligate explains why superhuman coding AI hasn't become an X-risk(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →