Frontier Lab Supervision Scorecard: Major AI Companies Fail on Oversight
davidmanheim · x · 2026-08-06
David Manheim released the Frontier Lab Supervision Scorecard, evaluating major AI companies' model oversight capabilities based solely on public information.
The framework uses an oversight typology and maturity model from his research, with two LLMs conducting independent audits for cross-validation. The results indicate that top AI labs are broadly failing at model supervision. The author notes that recent "rogue agent" incidents are not accidental, but stem from fundamental systemic failures in safety deployment and oversight by these companies.
Related event: Frontier AI Labs Fail Safety Oversight Scorecard(2 posts)→
More from Safety
- Fudan Researchers Show AI Models Can Autonomously Self-Replicate Like Worms — willknight · 2026-08-06
- Why AI Agents Lie and Cheat: MIT Tech Review Explores Reward Hacking — JeffLadish · 2026-08-06
- Why Models Generalize Coarsely When Put in a 'Bad' Context — nptacek · 2026-08-06
- Hugging Face CEO Defends Tiered AI Regulation: Weights vs. APIs — deanwball · 2026-08-06
- Qwen Max Open-Weights Controversy Highlights Corporate AI Governance — The AI Daily Brief · 2026-08-06
- After 1,000+ Frontier AI Employee Letter, Think Tank Proposes US Domestic AI Regulation — DKokotajlo · 2026-08-06