Ex-OpenAI alignment researcher slams Alexandr Wang's alignment pledge: employee-level audits or bust
austinc3301 · x · 2026-09-13
Former OpenAI alignment researcher Micah Carroll publicly called out Meta Superintelligence Labs chief Alexandr Wang's alignment remarks as 'weak sauce.'
- Wang had said MSL is rapidly scaling alignment efforts and believes alignment may become the gating factor for scaling near the frontier.
- Carroll countered that only employee-like third-party auditing counts, contrasting with Dario's pledge to grant third-party evaluators employee-level access.
- The spat highlights the emerging battle over safety commitments: self-assessment vs independent audit.
More from AGI Musings
- Alignment Is an Algorithm Problem: Why RL Can't Optimize "Don't Harm Humans" — MillionInt · 2026-09-14
- Chilson recommends Hazlett's 'The Political Spectrum' on Bell monopoly and innovation — neil_chilson · 2026-09-14
- Raphael Millière: under uncertainty, act on consensus AI mitigations now, debate later — raphaelmilliere · 2026-09-14
- Ilya Sutskever sparks backlash: 'If I die to killer AI, I want it to be American' — bennash · 2026-09-14
- "They Have No Moat and Everybody Is About to Find Out" — ayushtweetshere · 2026-09-14
- Daniel Lemire: AI may force a reshaping of academic hierarchies and usher in a new intellectual golden age — lemire · 2026-09-14