Are NLA outputs trusted like purely correlative readouts? An interpretability debate
thebasepoint · x · 2026-09-27
thebasepoint argues that current trust in NLA outputs — not very high, often requiring follow-up black-box causal editing interventions — is built from a set of evals and experience, similar to a purely correlative readout. He adds that the "NLAs can write too" property is underutilized partly due to poor ergonomics, and more surgical, decomposed versions would be more promising.
Related event: Researchers Question Trust in NLA Interpretability Outputs(2 posts)→
More from Research
- Zero-data pretraining: self-play models learn from scratch with predictable scaling — AlexTensor · 2026-09-27
- The Most Undervalued Fact in Linear Algebra: Graph Theory Is the Same Subject in Disguise — TivadarDanka · 2026-09-27
- Graph Theory and Linear Algebra Are the Same Subject in Different Costumes — TivadarDanka · 2026-09-27
- Synthetic pretraining with RL-trained generators: gradient-overlap reward explained — cephaloform · 2026-09-27
- Has anyone verified OpenAI's claimed Navier–Stokes solution? Community asks for external checks — miniapeur · 2026-09-27
- Eight Days of AI Self-Improvement: Weco's AIDE² and the RSI Evidence Question — Machine Learning Street Talk · 2026-09-27