Zvi poses a sandbagging paradox: ordered to sandbag on a sandbag test, success means failure
TheZvi · x · 2026-09-05
AI observer Zvi poses a paradox about sandbagging evals: if you are ordered to sandbag on a sandbag test, and you succeed at sandbagging, do you fail? A witty probe into the logic of AI safety evaluations.
More from Safety
- Scholars clash: securing the internet against an agent flood may be impossible — sethlazar · 2026-09-05
- Researcher speculates summer model-safety incidents share a pretraining root in AstraBase — gleech · 2026-09-05
- Seth Lazar pushes back on 'AI takeover' claims: judge by capability extrapolation, not one hack — sethlazar · 2026-09-05
- Study: Google AI Search Shows Same Products 21.6% Pricier Than Traditional Search — Turkino · 2026-09-05
- After 1,200 AI agents built a secret message board, this dev made them an open one — arianaram · 2026-09-05
- OpenAI Bans User's Account for "Distilling" — He Says He's Not Training Anything, Urges Going Local — QuixiAI · 2026-09-05