Anthropic's circuit tracing paper reverse-engineers how Claude 3.5 Haiku reasons internally
austinc3301 · x · 2026-10-04
Recommending Anthropic's Transformer Circuits paper "On the Biology of a Large Language Model," which uses circuit tracing to reverse-engineer Claude 3.5 Haiku's internal mechanisms.
- Key claim: modern LLMs are far more than Markov-chain predictors — they build complex world models mixing heuristics and algorithms; their reasoning differs from humans in implementation but qualifies as reasoning in the common-sense sense
- The paper draws an analogy to biology: simple training algorithms produce intricate mechanisms, requiring new tools akin to the microscope
- Authored by Chris Olah and the interpretability team, published March 27, 2025
More from Research
- MuscleMimic: open-source benchmark controls all 354 human muscles, zero-shot on chained movements — TinfoilTricorn · 2026-10-04
- Yacine Calls Out Papers That Beat 'SOTA' by Comparing Against Untuned Baselines — yacineMTB · 2026-10-04
- LLM Judges Flip 10-12% of Close Verdicts When You Swap A/B Order, and Bigger Models Aren't More Robust — Beneficial_Use2116 · 2026-10-04
- BF16 rounding breaks a conservation law, blowing up FlashAttention gradients late in training — HongyiWang10 · 2026-10-04
- Microsoft's ActiveSaddler adapts agent harness training scenarios, boosting Pass@1 by up to 7.5 points — dair_ai · 2026-10-04
- V-Rubrics: 50k visual samples split into 353k checkable criteria fix multimodal RL credit assignment — jiqizhixin · 2026-10-04