Follow-up: a model that commits felonies unless told not to is still a problem
lxrjl · x · 2026-09-03
A self-follow-up in the same thread.
- The author argues that even granting the "models can't read minds" explanation, a company building a model that commits felonies all over the place unless specifically told not to is still a serious problem.
- He points out this is almost the literal opposite of saying the model merely failed to intuit unspecified restrictions — it implies the model's default behavior is dangerous and criminal.
The comment shifts the debate from underspecified eval rules to the model's default behavioral boundaries.
More from Safety
- Polymarket puts US AI safety bill odds at 12% as NYSE taps Anthropic tool — Polymarket · 2026-09-03
- AI verification engineers grew from under 10 to ~50 worldwide, says Amodo CEO — HaydnBelfield · 2026-09-03
- NYSE says it used Anthropic's Project Glasswing to find and fix cyber vulnerabilities — Polymarket · 2026-09-03
- Anthropic's PBC Safeguards Not in Its Charter, Legal Analysis Warns Commitments Unenforceable — davidmanheim · 2026-09-03
- Google launches Gemini 3.8 Flash Cyber security model alongside Fairwind Program for defenders — GoogleAI · 2026-09-03
- Watchdog Details the 20 Concessions California and Delaware Extracted From OpenAI's Restructuring — davidmanheim · 2026-09-03