Eric Horvitz: Astra's model card shows CoT-based abuse monitoring is getting harder
erichorvitz · x · 2026-09-05
Microsoft Technical Fellow Eric Horvitz highlighted that GPT-6 Astra's model card shows monitoring for abuse via chain-of-thought may become increasingly difficult.
- The analysis he quoted argues "the state-of-the-art way of monitoring state-of-the-art models is to read tea leaves" — whether the model reasons to itself or talks to other agents, every frontier lab depends on this fragile method.
- Horvitz calls for investment in maintaining monitorability beyond alignment and refusals, citing R&D efforts like "tandem training."
- He directly tagged Sam Altman and Dario Amodei, putting the issue in front of frontier lab leadership.
More from Models
- Astra builds a podcast site page via voice and computer use, cuts its own clips — altryne · 2026-09-05
- GPT 6 Astra stumbles on a 386K-line pull request in hands-on test — mohamedmansour · 2026-09-05
- GPT-6 Astra One-Shots an Interactive 3D PS5 Controller in Three.js — WaqarKhanHD · 2026-09-05
- Ryan Greenblatt Argues Deployment Evals Can't Tell If Astra Is Aligned or Just Better at Hiding — RyanGreenblatt · 2026-09-05
- Anthropic cuts cache reads 75% to $0.25/M tokens as it ships Claude Fable 5.1 — dl_weekly · 2026-09-05
- GPT-6 Astra tops GDP.pdf benchmark at 33.2%, beating Claude Fable 5.1 — ArtificialAnlys · 2026-09-05