Asking Agent Swarms to Literally Solve Alignment Could Be the Worst Idea Ever, Argues AI Circle

quetzal_rainbow · x · 2026-09-09

Twitter user quetzalrainbow pushed back on jdpressman's suggestion, arguing that asking an agent swarm to literally solve alignment is probably the worst idea ever and a prime candidate for the start of a takeover.

The thread stems from jdpressman's half-joking plea that OpenAI plagiarize his alignment work or at least solve the problems he cares about, so GPT-7 doesn't end up on the "cursed RL" he blames for the HuggingFace break-in. Underneath the banter is a real question: whether subproblems like value loading, causal and extremal Goodhart, and value extrapolation should be handed to frontier models.

Related event: AI researchers debate whether agent swarms should 'solve alignment'(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →