Open interpretability problem: convincingly editing a model's in-context beliefs

thebasepoint · x · 2026-09-27

In a discussion with researcher Jack W. Lindsey, thebasepoint highlights one open problem among many in interpretability: is it possible to convincingly edit a model's in-context belief about some state of the world, or its current goal?

He argues this capability is crucial for running good causal experiments on models, and it remains essentially unsolved.

Original post →

More from Research

Research channel →