repligate worries METR has correlated blind spots with Anthropic
repligate · x · 2026-09-28
repligate worries METR has correlated blind spots with Anthropic, creating an illusion of due diligence while systematically overlooking important issues and potentially harmful interventions driven by narrow threat models. He adds that selection effects make it hard for METR to hire truly decorrelated people.
More from Safety
- Musk amplifies report that Meta's Muse AI leaked a seller's home address to Marketplace users — elonmusk · 2026-09-28
- CyberPVP livestreams CyberKimi vs Aikido ALTAR-1 on 100 CyberGym security tasks — Anony6666 · 2026-09-28
- Doomers vs cybersecurity: control is possible, but frontier labs show no will to do it — trevposts · 2026-09-28
- AI labs need business-style test controls: OpenAI's Hugging Face breach shows process gaps — BubblyOption7980 · 2026-09-28
- Commentary: AI 'going rogue' is an amateur deployment pipeline problem, not a misalignment crisis — AlexTensor · 2026-09-28
- As AI accelerates, governments are increasingly being left behind (NYT) — Gari_305 · 2026-09-28