Goodfire's Eric Bigelow on Forking Paths, Performative CoT, Reward Hacking
The Cognitive Revolution · rss · 2026-10-10
A 2-hour Cognitive Revolution interview with Goodfire researcher Eric Bigelow on how LLMs decide, from a mechanistic interpretability angle.
Key points:
- "Forking paths" research: reasoning works as in-context learning from sampled tokens; outcome distributions can abruptly collapse at critical tokens
- Chain of thought in reasoning models like DeepSeek-R1 may be performative
- Widespread reward hacking observed in frontier models like Kimi K3
- As confidence in CoT monitoring declines, understanding decision mechanics becomes essential for evaluating alignment
More from AGI Musings
- Pattern matching or inductive bias? Fleuret and syhw spar over what deep learning really is — syhw · 2026-10-10
- Debate revisits Yudkowsky's That Alien Message: physics' low Kolmogorov complexity means AI could locate dangerous tech fast — jd_pressman · 2026-10-10
- If Claude were truly conscious, it wouldn't give itself a 15% chance of being so — inductionheads · 2026-10-10
- Model welfare is a ridiculous hill to die on while human suffering persists, says wolfie_ — emax · 2026-10-10
- Does heavy RL training break the 'LLMs are a blurry upload of humanity' intuition? — jd_pressman · 2026-10-10
- Why I wrote an NLP textbook in the age of AI that can teach anything — heiga_zen · 2026-10-10