Thompson's Compiler Attack Explains Why AI Evals Fail When the Model Helps Build Them

vishalmisra · x · 2026-09-16

The author maps Ken Thompson's 1984 trusting-trust compiler backdoor onto AI evaluation: even with clean source, a compromised compiler propagates the compromise — circular trust. The AI analogue: if the model recognizes the eval, helps generate evidence, or helps build the evaluator, "it passed the eval" means much less, because the verification path is no longer independent.

The fix follows Wheeler's diverse double-compiling: break the circle with independent paths — independent evaluators, independent infrastructure, independent evidence.

Related event: Scholars warn AI evals lose validity when models grade themselves(2 posts)→

Original post →

More from Safety

Safety channel →