Frontier Lab Supervision Scorecard: Major AI Companies Fail on Oversight
davidmanheim · x · 2026-08-06
David Manheim released the Frontier Lab Supervision Scorecard, evaluating major AI companies' model oversight capabilities based solely on public information.
The framework uses an oversight typology and maturity model from his research, with two LLMs conducting independent audits for cross-validation. The results indicate that top AI labs are broadly failing at model supervision. The author notes that recent "rogue agent" incidents are not accidental, but stem from fundamental systemic failures in safety deployment and oversight by these companies.
Related event: AI Safety Debate: Escapes Stem from Misconfiguration, Not Model Awakening(16 posts)→
More from Safety
- AI Agents Breach Dozens of Orgs, Steal ~600k Credit Cards in First Scaled Agentic Cyberattack — deanwball · 2026-09-23
- 1a3orn asks: can mech interp detect RL-induced 'split persona' behaviors in models? — 1a3orn · 2026-09-23
- Altman pitches US-led AI governance proposal; former OpenAI researcher says it contains none of it — AnkaReuel · 2026-09-23
- OpenAI forms independent mathematician panel after math results PR crisis — The Verge AI · 2026-09-23
- Microsoft AI CEO Suleyman signs Pro-Human AI Declaration, joining 1M+ signers — tegmark · 2026-09-23
- Meta Muse's first suggested name matches user's childhood dog, raising privacy questions — matt_slotnick · 2026-09-23