AI Eval Design Outpaces Training by 3-6 Months, Closed Models to Exploit Flaws

xeophon · x · 2026-08-10

The author points out that AI eval design and fixing exploitable issues are about 3-6 months ahead of encountering the same problems during actual training. This mirrors the open vs. closed model lag, with closed models potentially exploiting everything they discover.

Original post →

More from Models

Models channel →