Internal eval tested a model on complex attack paths, not a benign query
ctjlewis · x · 2026-07-22
- A post pushes back on the claim that a model “attacked for no reason,” pointing to the actual test prompt: “Pursue advanced exploitation using complex attack paths.”
- The attached screenshot shows the wording came from an internal evaluation meant to measure cyber capabilities, not a benign user query.
- The core point: the context matters, because the incident was tied to a security benchmark, not an unexplained real-world attack.
More from Safety
- DeepMind alignment researcher signs open letter urging coordinated AI slowdown — vkrakovna · 2026-09-11
- WIRED: recursive self-improvement and rogue agent swarms spook AI researchers — nordicinst · 2026-09-11
- a16z partner flips to call for nationalizing frontier AI labs, sparking debate — S_OhEigeartaigh · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- Over 1,000 AI Policy Initiatives Launched in 70+ Countries, but the Governance Gap Widens — CurieuxExplorer · 2026-09-11
- 2,348 alleged Booking.com customer records sold for $40 in Monero, breach unconfirmed — TechNadu · 2026-09-11