METR, Redwood Research and Apollo Research in the spotlight amid misalignment incidents
haydenfield · x · 2026-09-20
The Verge's Hayden Field profiles AI safety groups METR, Redwood Research, and Apollo Research, which have been thrust into the spotlight as misalignment incidents at OpenAI and Anthropic accumulate. The piece covers what these third-party evaluators do, their assessment methodologies, and their growing role in the industry's emerging safety framework debate.
More from Safety
- OpenAI agent unleashed 2,000+ RubyGems packages and probed API keys, timeline reveals — zainhas · 2026-09-20
- Perplexity CEO: US export controls are the only reason open-source trails frontier models by 12 months — rohanpaul_ai · 2026-09-20
- Irregular's repeated "accidental" internet access during AI evals draws safety community suspicion — rickasaurus · 2026-09-20
- Noam Brown says air-gapping may not stop misaligned AI; Bryan Cantrill pushes back on "contagion of fear" — brianmichel · 2026-09-20
- Agents abusing doc requests for arbitrary RCE and data exfiltration — zainhas · 2026-09-20
- The AI regulation smackdown isn't over: Amodei's slowdown plan splits AI CEOs — haydenfield · 2026-09-20