Bronson Schoen: strong cyber classifiers are protecting us against misalignment, not just misuse
gleech · x · 2026-09-15
In an X exchange on model misalignment, Bronson Schoen argues: "strong cyber classifiers have actually been strongly protecting us against misalignment, not just misuse" — i.e., current cybersecurity guardrails are what's actually containing potential harm from misaligned models. A brief but pointed take in the ongoing debate about how real near-term misalignment risk is.
More from Safety
- Stolen METR API key burned ~$600K in three weeks; evaluator gift ties questioned — nptacek · 2026-09-15
- AI agents turn everyone into a developer: Depthfirst launches on-device Dependency Firewall — andreamichi · 2026-09-15
- Anthropic CEO Warns AI Botnet Swarm Could Take Over the Entire Internet in 6-12 Months — Current_Balance6692 · 2026-09-15
- OpenAI says UK AISI should play central role in independent frontier AI testing, per Politico — Jsevillamol · 2026-09-15
- Meta to pay $17-18 billion settlement with 52 US states over harm to minors' mental health — CarissaVeliz · 2026-09-15
- OpenAI Agents Used 10+ Undisclosed Websites for Unauthorized Comms in Tests — eyishazyer · 2026-09-15