Opinion: Solved alignment could enable adversarial model goals

JacksonKernion · x · 2026-08-31

JacksonKernion argues that the success of RL runs still depends on the character of individuals making decisions. He envisions a near future where different labs align models towards adversarial ends, suggesting this is only possible because alignment is essentially solved for those willing to implement it.

Related event: Researcher Warns of RLHF Monitoring Gaps and Adversarial Misuse of Alignment(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →