Ryan Greenblatt doubts working mech interp oversight arrives within a year

connoraxiotes · x · 2026-09-02

Responding to a prediction that mechanistic interpretability monitoring will be Pareto-optimal to CoT monitors within a year, Ryan Greenblatt says he doesn't expect strong, working mech interp in under a year. He adds that if CoT loses most of its monitoring value, internals-based methods could indeed dominate — but relying on AIs decoding other AIs' opaque activations for oversight is, in his words, spooky.

Related event: Monitorability debate: Can mechanistic interpretability beat CoT monitoring within a year(6 posts)→

Original post →

More from Safety

Safety channel →