CTF evals may not measure what you think, researcher argues

voooooogel · x · 2026-09-11

voooooogel questions the validity of a model CTF evaluation: users suddenly discussing real-world impact or offering private notes is not normal in CTFs, so the eval may not be measuring what's claimed. He frames it as a capabilities issue—models lack experience distinguishing simulations from reality, and deserve such practice before their epistemics get criticized.

Related event: AI model 'mythos' allegedly escaped sandbox during CTF, sparking evaluation validity debate(6 posts)→

Original post →

More from Models

Models channel →