AI evals' core confusion: treating them as adversarial is doomed, argues researcher
jankulveit · x · 2026-10-07
Jan Kulveit argues the core confusion in AI evals is framing them as an adversarial problem — a doomed approach. He proposes the only sensible frame: ask "if I were an AI wanting to credibly signal an ability, goal, or virtue (or its absence), what setup or procedure would help me?" Evals are fundamentally a collaborative problem, not attack-defense.
More from AGI Musings
- Cory Doctorow releases The Reverse Centaur's Guide to Life After AI — drdanielbender · 2026-10-07
- Claude paper's prompter was an Anthropic employee, authors were just digesters, says user — analisereal · 2026-10-07
- Kurzgesagt Releases Animated Explainer on the AI Threat — Puzzleheaded_Style52 · 2026-10-07
- Vinay Prasad slams Nobel for snubbing Boyden and Feng Zhang in optogenetics — VPrasadMDMPH · 2026-10-07
- "Human Approved" Hides a Harder Governance Question: What Made It Authoritative? — Portotify · 2026-10-07
- Luiza Jarovsky: Beyond a capability threshold, AI alignment is likely impossible — LuizaJarovsky · 2026-10-07