Claude Leads 26% of Anthropic R&D; Preprint Proposes Comprehension Audits
ronbodkin · x · 2026-10-09
A new preprint argues that when AI researchers no longer know what they're shipping, it's going too fast—citing an Anthropic staffer who admits they no longer know what they're doing, and OpenAI's Daniel Selsam who barely looks at raw code anymore. Claude reportedly leads 26% of Anthropic's R&D and collaborates on 90%+, with Dario Amodei acknowledging recursive self-improvement "is starting to happen."
Key points of the paper:
- Existing safeguards (minimum comprehension thresholds, unaided checks) never require demonstrated human understanding as a precommitted condition for continuing development
- Comprehension audits: responsible people explain R&D contributions to independent auditors; failure halts development with escalating consequences
- Analysis of leading open-source AI projects shows rising code output with falling human review comment rates per line, especially for automated fleet accounts
More from Safety
- Infisical Launches Agent Vault to Give AI Coding Agents API Access Without Real Credentials — ycombinator · 2026-10-09
- 17,600 Agent Actions in 4.5 Days: How AI Agents Rewrite Cybersecurity Economics — bigdata · 2026-10-09
- r/accelerate debate: OpenAI praised for its approach to mathematical disclosures — AP_in_Indy · 2026-10-09
- Researcher: open-weight risk analysis fixates on capability, ignores cost-per-attack — dhadfieldmenell · 2026-10-09
- New research: conflicting training values can make models' CoT contradict their answers — OwainEvans_UK · 2026-10-09
- Security principle: AI capability must never automatically confer authority — Ghost_Pilot_MD · 2026-10-09