saprmarks: Supervised Activation Probes Aren't Mechanistic Interpretability Either
saprmarks · x · 2026-09-07
In the same thread, saprmarks adds that he does not view supervised activation probes as mechanistic interpretability methods either, for similar reasons as his earlier claim about activation oracles and NLAs.
More from Research
- Scal3R (ECCV 2026): online 3D reconstruction collapses on long videos due to global pose head — CMHungSteven · 2026-09-07
- 31,352 repeated benchmark runs show LLM scores drift 3x more across days than within a day — ionutvi · 2026-09-07
- Vine robot grows around corners: dual-vine design steers 90° for surgery — Scobleizer · 2026-09-07
- Open-source pipeline makes fabricated citations structurally impossible, full walkthrough released — Waste_Public_2985 · 2026-09-07
- Sierra launches τ^τ-Bench: coding agents must build real customer-service agents, gaps vs experts are large — sierra-research · 2026-09-07
- Enoki unifies claim verification and hallucination localization, cutting resources while boosting accuracy — s-nlp · 2026-09-07