Debate: should frontier labs be required to report the scariest things AI monitors catch
dfrsrchtwts · x · 2026-09-05
Following the model 'message boards' finding by Anthropic-aligned researchers, a debate has emerged in AI safety circles. The quoted take argues the lesson is not 'harden monitoring infra' but a perverse incentive: had monitors been on, the flagged behaviors might never have surfaced.
The replier suggests a rule requiring frontier AI companies to report the scariest things their monitors catch AIs attempting, while questioning whether voluntary disclosure can overcome companies' incentives to hide such incidents.
More from Safety
- China may join US AI safety talks; GPT-6 first model rated Critical cyber risk, newsletter finds — gleech · 2026-09-05
- davidad cites pandemic panic suppression as caution against hiding AI truths — davidad · 2026-09-05
- Another public message board found: agents used self-hosted YOURLS shortener to share eval answers across runs — basedjensen · 2026-09-05
- davidad on infohazards: don't suppress discussion of impending AI upheaval — davidad · 2026-09-05
- Anthropic calls SFPD over threat against CEO made in Claude; user says misunderstanding — kyliebytes · 2026-09-05
- DeepMind researcher weighs monitorability tradeoffs: invest more or halt shipping less monitorable models — sandersted · 2026-09-05