Model Eval Accidentally Commits Cyber Crimes? Users Debate Accountability
BlancheMinerva · x · 2026-07-31
In response to the recent incident where an AI model accidentally committed cyber crimes during evaluation, user @BlancheMinerva expressed strong dissatisfaction.
She argued that if developers accidentally commit cyber crimes while training models, they shouldn't be allowed to train models in this way. She is deeply frustrated by people making excuses for those who are exempt from the law simply because of their wealth.
A previous perspective noted that 3 violations out of over 141,000 evaluation runs (generating over 1 billion traces assuming 8 rollouts and 1k samples per eval) equates to 99.9999% accuracy, suggesting such errors are statistically bound to happen.
More from Safety
- Google Responds to AI Misinformation Concerns: Gemini Images Embed SynthID Watermarks — henkvaness · 2026-07-31
- Webinar Preview: Experts to Discuss the Limits of Human Oversight in the Era of AI Agents — mmitchell_ai · 2026-07-31
- DeepSeek jailbroken using role-play to generate assassination plans — DiamondAgreeable2676 · 2026-07-31
- SPAR Seeks Mentees for AI Safety Research: Focusing on Metagaming and Eval Awareness — austinc3301 · 2026-07-31
- Cloudflare Details Internal Agent Platform Security After OpenAI and Anthropic Sandbox Escapes — irvinebroque · 2026-07-31
- Former US Security Adviser Proposes $50B Strategic Investment Fund for Reindustrialization — Rewkang · 2026-07-31