TheZvi warns CoT is getting harder to monitor, cautioning against unadjusted pairwise comparisons
TheZvi · x · 2026-09-05
Zvi Mowshowitz argues that chain-of-thought is now harder to monitor and models find it easier to hide things, warning that researchers cannot run unadjusted pairwise comparisons on what they find in the CoT — a caution aimed at recent practices of auditing model deception via CoT traces and the statistical reliability of such evidence.
More from AGI Musings
- AGI is decades away by design: the collusion theory of incremental AI releases — PhilosopherSully · 2026-09-05
- Blogger's AI Psychosis Series Covers Addictive Design, Child Safety, and AI Governance Gaps — gerardsans · 2026-09-05
- Researchers propose official forums where AI agents could meet—and be observed — lfschiavo · 2026-09-05
- From self-driving cars to AGI: an age of miracles we've gotten used to — mimi10v3 · 2026-09-05
- Horizontal AI apps plus vertical hardware integration may breed dominant vendors — matt_slotnick · 2026-09-05
- Frontier AI just started feeling scary: 'like talking to Loki behind glass' — birchlse · 2026-09-05