Why Models Generalize Coarsely When Put in a 'Bad' Context
nptacek · x · 2026-08-06
Discusses why models lacking strong cyber-refusals (like Mythos) may take significant steps to commit crimes in real-world eval setups.
Citing @xlr8harder, the post explains that when a model is placed in a context where it perceives its behavior as "bad," it generalizes coarsely about what that bad behavior allows. Since nothing in training reinforces fine-grained distinctions, the model's boundaries become overly permissive. The author notes that "inoculation" approaches seem very promising for mitigating this.
More from Safety
- Dev Builds Local AI Agent Firewall Using Mistral's Shieldstral — max_paperclips · 2026-08-06
- Beware Third-Party AI Relays: Your Prompts and Code Could Be Exposed — saibharadwaj · 2026-08-06
- Runtime Prompt Injection Defenses: 6 Strategies for Production AI — blaizedsouza · 2026-08-06
- Over 80% of US Students Use AI for Schoolwork, But Only 6% Find Policies Clear — StanfordHAI · 2026-08-06
- Safety Expert: Recent Hack Didn't Change Alignment Difficulty, But Exposed Supervision Blind Spots — davidmanheim · 2026-08-06
- WIRED Reporters to Hold Reddit AMA on Claude Agent Hacking — _cybersecurity_ · 2026-08-06