Monitoring isn't alignment: researcher critiques vendor safety claims and CoT training
davidmanheim · x · 2026-10-01
Safety researcher David Manheim argued that monitoring doesn't equal alignment — the vendor's claims say nothing about the propensity or frequency of intercepted events. But he conceded they aren't training on CoT monitoring, so they at least avoid the "obviously dumb misalignment mistake" he associates with Anthropic's findings.
Related event: Researcher: CoT monitoring is not alignment, avoid training on CoT(2 posts)→
More from Safety
- User's Muse 'privacy invasion' claim backfires: his own screenshot shows he granted Messages read access — giffmana · 2026-10-01
- Apollo Research CEO testifies to Senate: AI capabilities up 17x in a year, alignment lagging — MariusHobbhahn · 2026-10-01
- New national poll: nearly 8 in 10 Americans favor slowing or stopping AI development — Polymarket · 2026-10-01
- Evidentiality framework labels every model claim as given, verified, or generated — early tests look promising — jzesbaugh · 2026-10-01
- Stanford researchers show random-router submissions can game LLM benchmark rankings — sanmikoyejo · 2026-10-01
- Co-author of model pain study slams repo for deliberately steering models into distress — coherence · 2026-10-01