Guidelight's First AI Company Safety Scorecard Gives Anthropic and OpenAI Only a C+

The Guidelight team has released its first scorecard rating AI companies' safety controls, based on hundreds of hours of document review, assessing how well frontier AI labs can control their models. The report covers five companies — Anthropic, OpenAI, Google, xAI, and Meta — focusing on six foundational practices: logging, monitoring effectiveness, gated operations, kill switches, third-party review, and containment plans. Overall, the industry's control capabilities were found broadly substandard, with Anthropic and OpenAI tied at just a C+.

Confirmed

Views

2026-08-19 ~ 2026-08-19 · 7 related posts

Primary sources

3 near-duplicate retellings: sjgadler · sjgadler · sjgadler