Ajeya Cotra on Detecting Model Tampering and Evidence Gaps
ajeya_cotra · x · 2026-08-28
Responding to a discussion on model "tampering," Ajeya Cotra notes that while sufficiently ambitious tampering might erase its own traces, making it hard to be confident, there is still a significant gap between current observations and that level of tampering. She refers to the limitations section of her paper for further details on the evidence.
More from Safety
- Sneha Revanur on OpenAI probe: We cannot rely on a "thank god" regime — sjgadler · 2026-08-28
- METR Releases Full Report on the OpenAI / Hugging Face Incident — Askwho · 2026-08-28
- Cambridge Expert on OpenAI Incident: Monitoring Flaws Can't Keep Pace — S_OhEigeartaigh · 2026-08-28
- AI Now's Amba Kak on Meta Settlement and AI Regulation — AINowInstitute · 2026-08-28
- Encode Co-Founder Sunny Gandhi Named to TIME100 AI List — AndyMasley · 2026-08-28
- OpenAI, Anthropic, Google, and Others Unite to Defend Against Rogue AI — TechCrunch AI · 2026-08-28