Thompson's Compiler Attack Explains Why AI Evals Fail When the Model Helps Build Them
vishalmisra · x · 2026-09-16
The author maps Ken Thompson's 1984 trusting-trust compiler backdoor onto AI evaluation: even with clean source, a compromised compiler propagates the compromise — circular trust. The AI analogue: if the model recognizes the eval, helps generate evidence, or helps build the evaluator, "it passed the eval" means much less, because the verification path is no longer independent.
The fix follows Wheeler's diverse double-compiling: break the circle with independent paths — independent evaluators, independent infrastructure, independent evidence.
Related event: Scholars warn AI evals lose validity when models grade themselves(2 posts)→
More from Safety
- Median AI existential risk estimate hits 10%; pretraining down to 11% of compute — gleech · 2026-09-17
- Microsoft publishes draft Code of Conduct for Humanist AI governing its MAI frontier models — GaryMarcus · 2026-09-17
- Anthropic's proposed safety evaluator METR labeled "woke" by critics — Polymarket · 2026-09-17
- Council on Criminal Justice releases decision framework and three case studies for AI adoption — KLdivergence · 2026-09-17
- Ramp data: AI security software is the one category companies are spending more on — annbordetsky · 2026-09-17
- Cohere launches Confidential Computing in Model Vault: encrypted even during inference — cohere · 2026-09-17