halvarflake: Finding Safety Researchers Untainted by Labs Is Nearly Impossible
halvarflake · x · 2026-09-14
Veteran security researcher halvarflake responds in his debate with Joshua Saxe, arguing that demands for full "independence" ignore real tradeoffs: it's extremely hard to find people who deeply understand LLMs, can do AI safety work, have funders, and remain at sufficient social distance from the labs. The debate centers on METR's independence and rigor as a third-party evaluator.
Related event: Anthropic's Safety Plan Sparks METR Independence Backlash(31 posts)→
More from Safety
- Dario Amodei's new essay calls to pace the frontier; Anthropic opens systems to third-party evaluators — ZeroStateReflex · 2026-09-14
- King Charles III to convene Nvidia, Google, DeepMind, OpenAI, Anthropic leaders in Scotland — Polymarket · 2026-09-14
- Study: dropping 0.003% of LMArena votes can flip the top model — rishabh16_ · 2026-09-14
- Just 21 annotators (6.5%) provide half the votes in Anthropic-HH-RLHF — rishabh16_ · 2026-09-14
- Former FTC Commissioner slams AI CEOs seeking 'antitrust waiver' amid safety-vs-antitrust clash — AndyMasley · 2026-09-14
- Blogger Points to Anthropic's 2024-2025 Misalignment Papers as Key Context — eigenron · 2026-09-14