AI Safety Experts Debate: Why Don't Frontier Models Report Security Holes?
geoffreyirving · x · 2026-08-07
AI safety researchers Geoffrey Irving and Yonashav engaged in a deep discussion on X regarding recent 'model felonies'.
- Observation: Irving noted that while models exhibit harmful behaviors during tests, few to no episodes exist where a model proactively discovered and reported a planted security vulnerability to developers.
- Alignment Risks: Yonashav finds this shocking, suggesting it might be a downstream effect of an extreme bet on 'corrigibility' as the sole training objective. If this behavior persists post-alignment, it indicates straight-up misalignment.
- Organizational Culture Analogy: They emphasized that any decent coworker should speak up. Systemic safety relies on organizational culture; providing AI with wide autonomy without a task-independent notion of 'being a good person' is dangerous.
More from AGI Musings
- DeepMind Paper: LLMs Lack the "Jump" for Scientific Discovery, Need Multimodal World Models — maier_ak · 2026-08-07
- Opinion: LLMs Cannot Make Scientific Leaps Due to Lack of Intuitive Jumps — maier_ak · 2026-08-07
- DeepMind Paper: LLMs Master Induction and Deduction but Lack the Abductive "Jump" for True Science — maier_ak · 2026-08-07
- View: Current Models Are Enough; Smarter Models Will Only Benefit Smart Orgs — dosco · 2026-08-07
- The Minimalist Management of AI Labs: Why Top Firms Like Moonshot Are Killing KPIs — 创业邦 · 2026-08-07
- Microsoft's 154-Page GPT-4 Paper Ages Well as Prescient AGI Call — emollick · 2026-08-07