Activation Oracles Author Weighs In: 'A Fundamentally Non-Mechanistic' Interp Technique
saprmarks · x · 2026-09-07
saprmarks, co-author of the activation oracles work, says on X that he does not consider activation oracles or NLAs to be mechanistic interpretability methods, pointing to a section in the blog post literally titled "Activation Oracles are a fundamentally non-mechanistic technique for interpreting LLM activations." The thread also debates whether probes deployed at frontier labs count as a mech interp win.
More from Research
- DeepMind-Princeton paper shows LLMs causally use confidence to decide whether to answer — GoogleDeepMind · 2026-09-07
- The attention triangle: diagnosing cross-modal semantic leakage in audio-video diffusion — tau · 2026-09-07
- Yandex researchers propose KV cache as an agent runtime, demo Qwen3.8 playing DOOM interactively — _puhsu · 2026-09-07
- ECCV 2026 workshop on deep learning era SfM set for Sept 9 with 3 speakers and 8 papers — ducha_aiki · 2026-09-07
- Actually queryable executables: a webserver whose binary and state are one SQLite file — bibryam · 2026-09-07
- PLANET lifts multi-object tracking into 3D scene geometry, hits SOTA on DanceTrack, SportsMOT & BFT — lealtaixe · 2026-09-07