Models trained with 'blindsight' toward their own CoT, researcher notes
voooooogel · x · 2026-09-13
In the thread on the airplane analogy, vooooogel argues we don't have access to human CoT either and make it work, but agrees much innovation has been strangled in the cradle. He highlights that models are now trained to have a weird blindsight toward their own chain-of-thought for the sake of anti-extraction — distorting model self-transparency by design.
Related event: 'Flight Instruments' Analogy Fuels Debate Over Hidden Chain-of-Thought(4 posts)→
More from AGI Musings
- Game theory of AI slowdown pledges: nobody slows down internally, everyone loses — haider1 · 2026-09-13
- The next SaaS shift: buying software that lets your agent work autonomously — eptwts · 2026-09-13
- AI solved a Millennium Prize Problem while he wrote about it: "We will never have normal times again" — voooooogel · 2026-09-13
- UK pilots three-week AI bootcamps in Preston to turn Neet youths into apprentices — nordicinst · 2026-09-13
- Video argues AI can't take everyone's jobs: excess demand means no mass layoffs — EGarrett28 · 2026-09-13
- Altman, Musk, and Hassabis back Amodei's call for independent AI oversight; OpenAI IPO delayed to 2027 — The Decoder · 2026-09-13