Interpretability researcher proposes predicting agent behavior by cross-referencing other agents' hidden states

jachiam0 · x · 2026-09-27

The author argues interpretability research should aim to reliably predict one agent's behavior by comparison to others, rather than requiring mechanistic explanations in a vacuum: cross-reference an agent's hidden states against hidden states of many other agents in comparable situations to form probability estimates of replicating known outcomes.

He suggests this approach could also help predict when two agents will align, ally, and cooperate — possibly another critical interpretability goal. The post ends by asking whether prior papers on this line of work already exist.

Original post →

More from AGI Musings

AGI Musings channel →