Jeff Ladish Questions Anthropic's Transparency and RSI Safety Commitments
AI safety researcher Jeff Ladish posted a flurry of comments in the same discussion thread on 09-07, centering on one pointed question: why has OpenAI recently outpaced Anthropic on transparency? He called on Anthropic to publish more detailed write-ups on both internal research acceleration (RSP-style topics) and international coordination, and said that given its past statements, he will hold transparency to an extremely high bar and speak plainly about what he has seen.
Confirmed
- Ladish expressed near-zero confidence in Anthropic's plans to constrain recursive self-improvement (RSI), and doesn't believe it would actually stop. By contrast, he noted OpenAI has claimed to have paused an RL training run for at least several weeks, while it's unclear whether Anthropic has ever done anything similar.
- He warned that "fully automated AI R&D" being treated as the default path deserves deep skepticism, saying it sounds like a disaster—possibly even literal human death.
- He said he doesn't want to see Anthropic fail and externalize massive risks onto the world, is happy to be proven wrong, and worries Anthropic has already fallen (or is falling) into the same risk gravity well as OpenAI and other labs.
- Citing Tim Hua's analysis, he commented that OpenAI was reportedly forced to bet on AI coding performance and sprint toward RSI because of Anthropic's strong momentum with Claude Code.
- He voiced direct concern about the AGI race: humans don't want to compete for resources with an optimization process far more powerful than themselves, and a piece of paper can't guarantee a share; the "can't slow innovation" logic holds for other technologies, but AI will catastrophically self-accelerate past a critical threshold (as he paraphrased, failing to slow the intelligence explosion could mean death).
- The discussion backdrop includes Anthropic's new article "When AI builds itself," which reports engineers' quarterly code output has grown 8x, plus Anthropic's August risk report (covered by Zvi Mowshowitz, generally positive but with reservations).
Why it matters
- This is a firsthand voice from a frontier-lab safety researcher publicly applying pressure, directly targeting the scope of Anthropic's disclosures and the credibility of its RSI pause mechanisms.
- The tension between extreme-risk narratives like fully automated AI R&D and intelligence explosions, and labs' actual development pace (e.g., Claude Code-driven acceleration), sits at the heart of today's AI governance debate.
2026-09-07 ~ 2026-09-07 · 10 related posts
Primary sources
- AI Safety Researcher Questions If OpenAI Is Now More Transparent Than Anthropic — JeffLadish ·
- Safety researcher Jeff Ladish: fully automated AI R&D sounds like 'catastrophe, plausibly literal death' — JeffLadish ·
- Jeff Ladish: little confidence Anthropic would ever stop an RSI run like OpenAI's claimed RL pause — JeffLadish ·
- Jeff Ladish: If we don't slow the intelligence explosion, we die — JeffLadish · 2026-09-07
- OpenAI reportedly pivoted to coding and RSI racing in response to Claude Code pressure — JeffLadish · 2026-09-07
- [source] AI Safety Researcher Questions If OpenAI Is Now More Transparent Than Anthropic — JeffLadish · 2026-09-07
- Jeff Ladish Follows Up: Wants Detailed Anthropic Write-ups on Research Acceleration and Coordination — JeffLadish · 2026-09-07
- Anthropic: Engineers Now Ship 8x More Code Per Quarter as AI Builds AI — JeffLadish · 2026-09-07
- Anthropic's August Risk Report Reveals New Model and Autonomy Risks — JeffLadish · 2026-09-07
- [source] Jeff Ladish: little confidence Anthropic would ever stop an RSI run like OpenAI's claimed RL pause — JeffLadish · 2026-09-07
- Jeff Ladish fears Anthropic is falling into the same attractors as other labs — JeffLadish · 2026-09-07
- [source] Safety researcher Jeff Ladish: fully automated AI R&D sounds like 'catastrophe, plausibly literal death' — JeffLadish · 2026-09-07
1 near-duplicate retellings: JeffLadish