Averaging OCR attention heads across layers accidentally yields a strong interpretability lens
benno_krojer · x · 2026-09-29
- bennokrojer shares an interpretability anecdote: the team wasn't originally looking for a general lens; sheridanfeucht was just trying to find OCR attention heads.
- Surprise finding: averaging those heads into one matrix — even across layers — makes a great lens.
- A reminder that general interpretability tools can emerge from very local exploration.
More from Research
- What do 56 AI benchmarks actually measure? COLM oral paper applies validity testing — ang3linawang · 2026-09-29
- Reasoning Models Stay Miscalibrated: EMNLP Studies on AI Confidence — zhaoran_wang · 2026-09-29
- Anthropic: Infrastructure Config Alone Swings Agentic Coding Benchmarks by 6 Points — giansegato · 2026-09-29
- GoFish demo: hierarchical chart construction lets you attach animations directly to chart structure — arvindsatya1 · 2026-09-29
- Eval author warns against over-indexing on benchmarks; Anthropic finds infra noise can swing scores 6 points — giansegato · 2026-09-29
- Claim: Transformers can now be pretrained with zeroth-order optimization, no backprop — teortaxesTex · 2026-09-29