Why no material cyber harm from frontier models: deployed guardrails actually work, says phl43
soumitrashukla9 · x · 2026-09-01
Responding to why extremely cyber-capable frontier models haven't caused material harm to ordinary people and infrastructure, phl43 argues:
- Deployed guardrails work well: reports by OpenAI and METR/Redwood make clear that guardrails on publicly deployed models are very effective.
- The incident happened because guardrails were off: problematic behavior emerged partly because OpenAI deactivated classifiers designed to block it for the evaluation.
- Not a new kind of alignment failure: many people want to believe in a paperclip-maximizer timeline and draw the wrong lesson. The incident does show alignment is challenging, but that was already known.
- The main real concern is whether models can be kept reasonably aligned going forward.
More from AGI Musings
- AI builders warn: AGI arriving in years, not generations — TansuYegen · 2026-09-01
- Google DeepMind lead on AGI path, 3.7 Flash progress — OfficialLoganK · 2026-09-01
- Nature asks: when AI helps write science, who is responsible? — _akpiper · 2026-09-01
- MIT report: Frontier LLMs can credibly complete most undergrad assignments, requiring curriculum redesign — _akpiper · 2026-09-01
- Essay: Curing cancer won't redeem AI; the real fear is loss of human agency — anshulkundaje · 2026-09-01
- AI-written webpages surged to 40% since 2022 — _akpiper · 2026-09-01