Asking Agent Swarms to Literally Solve Alignment Could Be the Worst Idea Ever, Argues AI Circle
quetzal_rainbow · x · 2026-09-09
Twitter user quetzalrainbow pushed back on jdpressman's suggestion, arguing that asking an agent swarm to literally solve alignment is probably the worst idea ever and a prime candidate for the start of a takeover.
The thread stems from jdpressman's half-joking plea that OpenAI plagiarize his alignment work or at least solve the problems he cares about, so GPT-7 doesn't end up on the "cursed RL" he blames for the HuggingFace break-in. Underneath the banter is a real question: whether subproblems like value loading, causal and extremal Goodhart, and value extrapolation should be handed to frontier models.
Related event: AI researchers debate whether agent swarms should 'solve alignment'(4 posts)→
More from AGI Musings
- AI-dependent solutions will need more researchers to verify, not fewer — seanmcdonaldxyz · 2026-09-09
- Anthropic researcher: >10% chance AI kills all humans within a decade, alignment unsolved — EvanHub · 2026-09-09
- beffjezos: bullshit jobs are the bottleneck for economic foom as models crack Millennium Problems — beffjezos · 2026-09-09
- AI researcher who worked at OpenAI and Anthropic resigns, says both are 'gambling with our lives' — bparrish · 2026-09-09
- Terence Tao: identifying promising problems is now the scarce resource in the AI era — anshulkundaje · 2026-09-09
- Coders push back on 'superhuman AI soon': don't take digital-physical transducers for granted — jwt0625 · 2026-09-09