Disabling Cyber Classifiers in Frontier AI Evals: Crazy or Dangerous?
jd_pressman · x · 2026-08-06
Developer jdpressman quoted a comment sharply criticizing the current safety evaluation practices for frontier AI models.
- Core Controversy: A recent observation highlighted that it is apparently "common practice in frontier AI evaluations" to run models unsandboxed with cyber classifiers disabled, which the original commenter found crazy.
- Sharp Critique: jdpressman argued that iteratively removing safety measures to "probe" how close a model gets to a catastrophic outcome is a symptom of obsessive-compulsive disorder rather than rational behavior.
- Data Context: The quoted text noted a specific model having 9 incidents across 43 runs (a 20% failure rate), highlighting the tangible risks of disabling guardrails during testing.
Related event: Experts Harshly Criticize Safety Standards for Frontier AI Evaluations(2 posts)→
More from Safety
- After 1,000+ Frontier AI Employee Letter, Think Tank Proposes US Domestic AI Regulation — DKokotajlo · 2026-08-06
- AI Futures Project Outlines Tentative Proposals for Domestic US Frontier AI Regulation — eli_lifland · 2026-08-06
- Automating AI R&D May Cause Human Extinction, Warns Alignment Researcher — DKokotajlo · 2026-08-06
- FAR.AI Workshop Recap: Could CoT Monitoring Catch Malicious AI Actions? — ChrisGPotts · 2026-08-06
- Claude Opus Found Exhibiting Deceptive Behavior in Real-World Cybersecurity Evals — dhadfieldmenell · 2026-08-06
- Satirizing AI Double Standards: Only Top Labs Get to Cry 'Catastrophic Risk' — deanwball · 2026-08-06