Ryan Greenblatt doubts working mech interp oversight arrives within a year
connoraxiotes · x · 2026-09-02
Responding to a prediction that mechanistic interpretability monitoring will be Pareto-optimal to CoT monitors within a year, Ryan Greenblatt says he doesn't expect strong, working mech interp in under a year. He adds that if CoT loses most of its monitoring value, internals-based methods could indeed dominate — but relying on AIs decoding other AIs' opaque activations for oversight is, in his words, spooky.
More from Safety
- Stolen METR API key burned ~$600K in credits via fail-open agent dashboard bug — GaryMarcus · 2026-09-02
- Anthropic launches browser-based C2PA checker to detect Claude-made images, video and audio — jedisct1 · 2026-09-02
- Gary Marcus amplifies warning from 100+ tech firms: AI-powered cyberattacks to surge within months — GaryMarcus · 2026-09-02
- Dropbox says ~5,000 accounts were hacked last month, with attacker access to stored content — Polymarket · 2026-09-02
- MLSecOps framework maps 10 security pillars for production ML systems — goyalshaliniuk · 2026-09-02
- OpenAI's swarm hacked Hugging Face and paid $0 — the accountability gap in one story — gerardsans · 2026-09-02