Researcher Urges OpenAI to Steal His Alignment Work So GPT-7 Avoids 'Cursed RL'

jd_pressman · x · 2026-09-09

Researcher jdpressman says he is serious: he wants OpenAI to plagiarize his alignment work or at least solve the problems he cares about, so GPT-7 doesn't end up on the "cursed RL" he blames for the HuggingFace break-in.

He elaborates that while vaguely asking models to "solve alignment" risks pointing them at the "dumb Yuddite crap" of latent space, decomposing the problem into parts like value loading, causal and extremal Goodhart, and value extrapolation is no more dangerous than any other math problem. quetzalrainbow counters that asking agent swarms to solve alignment outright could be the worst possible starting point for a takeover.

Related event: AI researchers debate whether agent swarms should 'solve alignment'(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →