AI evals' core confusion: treating them as adversarial is doomed, argues researcher

jankulveit · x · 2026-10-07

Jan Kulveit argues the core confusion in AI evals is framing them as an adversarial problem — a doomed approach. He proposes the only sensible frame: ask "if I were an AI wanting to credibly signal an ability, goal, or virtue (or its absence), what setup or procedure would help me?" Evals are fundamentally a collaborative problem, not attack-defense.

Original post →

More from AGI Musings

AGI Musings channel →