Frontier AI Safety Lags: OpenAI and Anthropic Score C+ in New Control Assessment
sjgadler · x · 2026-08-19
Guidelight released its first scorecard assessing the control practices of frontier AI companies (Anthropic, OpenAI, Google, xAI, Meta). The evaluation focuses on six foundational practices including logging, monitoring efficacy, gated actions, circuit breaking, third-party review, and containment plans. The findings reveal that basic control practices are, at best, partially implemented across all companies, with the highest scores being a C+.
More from Safety
- AI Summer Trends: Multi-Agent Systems and Chain-of-Thought — nptacek · 2026-08-19
- Real AI risks lie outside the model: permissions, data, and presentation — bigdata · 2026-08-19
- Deepfakes accounted for 52% of major AI incidents in 2025 — Comfortable_Gene5180 · 2026-08-19
- Musk retweets discussion on AI oligopoly alignment — zacharylipton · 2026-08-19
- Zvi on Watermarks and Constraints: Not Primarily x-risk Reduction — TheZvi · 2026-08-19
- PA Governor Enacts Nation's Strictest AI Data Center Standards via Executive Order — TinfoilTricorn · 2026-08-19