Interp researcher pushes back: probes have advanced well beyond pre-LLM-era techniques
aryaman2020 · x · 2026-09-07
In a debate over the value of probes, the author rebuts the claim that probes are simple and unimproved by mechanistic interpretability research:
- Frontier labs deploying probes in production largely rely on interpretability researchers, who should claim probes as a big interp win
- Cites two academic examples where causal interp advances improved probes and vice versa
- Probe architectures (attention probes, whole transformer probes) and hyperparameter tuning are now very different from the pre-LLM era
- Asks: if not the post-LLM wave of interp research, who do you credit—and who do you fund to work on this?
More from AGI Musings
- Predicting a Lean community schism: human-readable proofs vs AI-generated slop — jacobaustin132 · 2026-09-07
- Astra's Stunning 3D Understanding Wows AI Execs, But Spatial Leaps Don't Equal Common Sense — GabGarrett · 2026-09-07
- After the Resy Bot Fiasco: x402 Payments Plus Proof-of-Human for Fairer Ticket Queues — kleffew94 · 2026-09-07
- Reddit Post Warns the AI Bubble Will Pop Brutally as Circular Compute Revenue Inflates Valuations — Richnix1551 · 2026-09-07
- Both Anthropic and OpenAI have won the AGI race, argues compute-contract analysis — teortaxesTex · 2026-09-07
- David Sacks: AI doom narratives already have serious money behind them, Anthropic IPO could enlarge it — rohanpaul_ai · 2026-09-07