Encoding What Models Think Is Easier Than It Seems, Says Researcher
joshua_saxe · x · 2026-09-06
Joshua Saxe argues the bar for encoding a model's reasoning is low: CoT quality is mediocre, and compressing aspects of model behavior into a compact representation from which one could reconstruct its 'thoughts' — akin to image captioning — seems achievable. The hard part, he says, is doing it in production and at scale, as much an engineering problem as a research one.
More from AGI Musings
- Narrative Agents: What Happens When AI Moves from Conversation to Relationship — begusgasper · 2026-09-06
- Blogger proposes RL paradigm that maximizes user agency and freedom — teortaxesTex · 2026-09-06
- Formalize All Human Math in a Year? Bold AI Plan Gets Eric Weinstein's Backing — AccBalanced · 2026-09-06
- Dev Overwhelmed by AI Release Blitz: GPT-6 Astra, Opus 5, Grok 4.6 and More — prasenx · 2026-09-06
- Ben Todd mocks AI risk debate: only focus on present dangers, never think ahead — ben_j_todd · 2026-09-06
- Dev after using fable and astra: 'Taste' is the highest abstraction in vibe coding — dosco · 2026-09-06