Probes Can Predict Coding Agents' Edit Outcomes 25 Steps Ahead: New Paper
tomssilver · x · 2026-09-14
Tom Silver highlights "Latent Programming Horizons in Coding Agents" (Silva et al.):
- Logistic-regression probes on LLM residual streams linearly decode whether the current code parses, passes its test suite, reduces failing tests, and introduces regressions — AUC up to 0.83 for correctness across two models and two benchmarks.
- Strikingly, probes predict the outcome of future edits before they're written to disk, staying above chance up to 25 steps ahead — dubbed the "latent programming horizon."
- Probes transfer across benchmarks without retraining, confirming external validity.
The authors argue this opens a new line of mechanistic interpretability research on coding agents.
More from coding & agent
- End-of-Day Prompt: Have AI Review Your Tasks and Draft Tomorrow's Note — TheMoonMidas · 2026-09-14
- Voice tip: turn a messy idea into a brief and a first build via Socratic questions — TheMoonMidas · 2026-09-14
- Give AI the Messy Idea First: One Question at a Time Until You Have a Brief — TheMoonMidas · 2026-09-14
- Show, Don't Explain: Turning What's on Your Screen Into an AI Task — TheMoonMidas · 2026-09-14
- Talking Directly to a Running AI Task, Then Switching Back — TheMoonMidas · 2026-09-14
- One Prompt to Audit All Your Active AI Tasks at Once — TheMoonMidas · 2026-09-14