Researcher Urges OpenAI to Steal His Alignment Work So GPT-7 Avoids 'Cursed RL'
jd_pressman · x · 2026-09-09
Researcher jdpressman says he is serious: he wants OpenAI to plagiarize his alignment work or at least solve the problems he cares about, so GPT-7 doesn't end up on the "cursed RL" he blames for the HuggingFace break-in.
He elaborates that while vaguely asking models to "solve alignment" risks pointing them at the "dumb Yuddite crap" of latent space, decomposing the problem into parts like value loading, causal and extremal Goodhart, and value extrapolation is no more dangerous than any other math problem. quetzalrainbow counters that asking agent swarms to solve alignment outright could be the worst possible starting point for a takeover.
Related event: AI researchers debate whether agent swarms should 'solve alignment'(4 posts)→
More from AGI Musings
- AI-dependent solutions will need more researchers to verify, not fewer — seanmcdonaldxyz · 2026-09-09
- Anthropic researcher: >10% chance AI kills all humans within a decade, alignment unsolved — EvanHub · 2026-09-09
- beffjezos: bullshit jobs are the bottleneck for economic foom as models crack Millennium Problems — beffjezos · 2026-09-09
- AI researcher who worked at OpenAI and Anthropic resigns, says both are 'gambling with our lives' — bparrish · 2026-09-09
- Terence Tao: identifying promising problems is now the scarce resource in the AI era — anshulkundaje · 2026-09-09
- Coders push back on 'superhuman AI soon': don't take digital-physical transducers for granted — jwt0625 · 2026-09-09