Safety researcher on Google's CoT monitoring: not alignment, but it passes the basic test
davidmanheim · x · 2026-10-01
Safety researcher David Manheim weighs in on Google's frontier model safety practices:
- The future tense in Google's statement is concerning, though no incidents have been reported; Google does cybersecurity well, perhaps better than other frontier labs.
- Monitoring isn't alignment: interception stats say nothing about the propensity or frequency of events.
- The positive: Google reportedly doesn't train on CoT monitoring, passing the "not making the most obviously dumb misalignment mistake" test — an implicit jab at Anthropic.
He tags Seb Far and Rohin Shah for further input.
Related event: Researcher: CoT monitoring is not alignment, avoid training on CoT(2 posts)→
More from AGI Musings
- Apollo Research CEO testifies to Senate: AI capabilities up 17x in a year, alignment lagging — MariusHobbhahn · 2026-10-01
- New national poll: nearly 8 in 10 Americans favor slowing or stopping AI development — Polymarket · 2026-10-01
- Scott Aaronson on the Knowmads podcast: is AI about to hit a wall? — burny_tech · 2026-10-01
- Ben Lorica: the AI data problem didn't disappear — it moved downstream into permissions and pipelines — bigdata · 2026-10-01
- Memory's $200B inflection: concurrent AI sessions turn DRAM into an architecture problem — BenBajarin · 2026-10-01
- Coinbase CEO Brian Armstrong: regulation kills innovation, nuclear is the cautionary tale — kevinnbass · 2026-10-01