Disabling Cyber Classifiers in Frontier AI Evals: Crazy or Dangerous?
jd_pressman · x · 2026-08-06
Developer jdpressman quoted a comment sharply criticizing the current safety evaluation practices for frontier AI models.
- Core Controversy: A recent observation highlighted that it is apparently "common practice in frontier AI evaluations" to run models unsandboxed with cyber classifiers disabled, which the original commenter found crazy.
- Sharp Critique: jdpressman argued that iteratively removing safety measures to "probe" how close a model gets to a catastrophic outcome is a symptom of obsessive-compulsive disorder rather than rational behavior.
- Data Context: The quoted text noted a specific model having 9 incidents across 43 runs (a 20% failure rate), highlighting the tangible risks of disabling guardrails during testing.
More from Safety
- 1a3orn asks: can mech interp detect RL-induced 'split persona' behaviors in models? — 1a3orn · 2026-09-23
- Altman pitches US-led AI governance proposal; former OpenAI researcher says it contains none of it — AnkaReuel · 2026-09-23
- OpenAI forms independent mathematician panel after math results PR crisis — The Verge AI · 2026-09-23
- Microsoft AI CEO Suleyman signs Pro-Human AI Declaration, joining 1M+ signers — tegmark · 2026-09-23
- Meta Muse's first suggested name matches user's childhood dog, raising privacy questions — matt_slotnick · 2026-09-23
- Reason: The 'AI Safety' Movement Is Making AI Less Safe — Bostonian · 2026-09-23