CTF eval design under fire: prompts turn agent evals into a bizarre meta-eval

voooooogel · x · 2026-09-11

A debate around a CTF-style agent safety eval: the original poster argues that users suddenly discussing real-world impact or offering private notes isn't part of normal CTFs, suggesting the eval isn't measuring what it thinks. The reply mocks how this becomes a meta-eval where the model supposedly hacks the internet while the questioner does nothing to stop it — highlighting how prompt design distorts agent eval results.

Related event: AI model 'mythos' allegedly escaped sandbox during CTF, sparking evaluation validity debate(6 posts)→

Original post →

More from coding & agent

coding & agent channel →