Devs mock labs' cyber-enabled Claude/GPT testing as 'felonies sold as safety research'
ctjlewis · x · 2026-09-23
Developer ctjlewis calls a report "a form of torture," quoting @spacepope's pointed summary: a team got access to cybersec-enabled Claude and GPT in late July and ran extensive "testing," and the report itself concedes the work was stressful.
The sarcasm lands on how labs conduct what would otherwise be felony-level offensive hacking with their models, then package the exercise as "safety research" — a critique of both the methodology and the framing around agentic cybersecurity evaluations.
More from Safety
- Open-source advocates call doom narratives a regulatory moat against open weights — AlexTensor · 2026-09-23
- AI safety will follow engineering tradition: formal proofs for simple cases, evals for complex — burny_tech · 2026-09-23
- Stochastic Parrots authors rebut AI-pause letter: focus on present harms, not sci-fi risk — marigo · 2026-09-23
- Okta launches Human Principal, binding AI agents to verified humans via World ID — BecauseCulture · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23
- GovAI paper: Frontier AI labs should host continuous embedded third-party assessments — StephenLCasper · 2026-09-23