Third-party evals won't speed alignment science; research transparency may work better
1a3orn · x · 2026-09-13
- In a debate with Daniel Kokotajlo, the author argues third-party evals only partially solve the principal-agent problem of "grading your own homework," but don't speed alignment science or enable coordination through common knowledge.
- Instead, investing in "research transparency" could (1) compress frontier labs' margins more, (2) yield 1-2 orders of magnitude more alignment science, and (3) be less susceptible to capture by labs.
More from Safety
- Satirical dialogue skewers Altman and Amodei for pushing 'safety' as a cartel — ziv_ravid · 2026-09-13
- Embedded AI Evaluators Need Double-Blind Evaluations for Real Credibility — Dr_Atoosa · 2026-09-13
- nic_carter: OpenAI sees tons of MNPI daily — an insider trading case is coming — AccBalanced · 2026-09-13
- Is compute AI's uranium? A nuclear analogy for frontier AI governance — geoffwolfe · 2026-09-13
- METR is now load-bearing, but US CAISI and UK AISI are absent from AI safety talks — teortaxesTex · 2026-09-13
- Cohere CEO Aidan Gomez: third-party AI audits are power projection, not real fixes — yacineMTB · 2026-09-13