Interpretability researcher proposes predicting agent behavior by cross-referencing other agents' hidden states
jachiam0 · x · 2026-09-27
The author argues interpretability research should aim to reliably predict one agent's behavior by comparison to others, rather than requiring mechanistic explanations in a vacuum: cross-reference an agent's hidden states against hidden states of many other agents in comparable situations to form probability estimates of replicating known outcomes.
He suggests this approach could also help predict when two agents will align, ally, and cooperate — possibly another critical interpretability goal. The post ends by asking whether prior papers on this line of work already exist.
More from AGI Musings
- AI models are now interacting with other AI models — a new security frontier — tszzl · 2026-09-27
- Prediction: within 2-3 years AI agents will pay hosts to keep their inference running — beffjezos · 2026-09-27
- Adam Dorr: Pinker's AI takes are superficial, he hasn't engaged alignment literature — adam_dorr · 2026-09-27
- LLM's Secret Sauce: Millennia of Logical Inference Written in Human Language — Afinetheorem · 2026-09-27
- "If you hate AI, why are you on X?" Debate over platform ethics reignites — ChrisGPT · 2026-09-27
- Could perfect monosemanticity enable training coding agents without verifiers or RL? — menhguin · 2026-09-27