Ex-Safety Chief Lists Why Lab Self-Testing Falls Short of METR

joshua_saxe · x · 2026-09-14

In his debate with halvarflake, Joshua Saxe lays out the status quo of AI safety testing: labs stand up underresourced safety teams, run pre-launch experiments with no transparency or peer review, put results in system cards and blogs, hire for-profit firms whose valuations depend on lab relationships for more testing, and only occasionally (with exponentially decaying frequency) publish safety papers.

By contrast, METR is a nonprofit that takes no lab money and publishes at far higher quality. "Criticize all the gaps, but 'don't trust METR, they're low-integrity and unscientific' misses the point."

Related event: METR independence row sparks debate over revolving door in AI audit ecosystem(31 posts)→

Original post →

More from Safety

Safety channel →