Why Models Generalize Coarsely When Put in a 'Bad' Context

nptacek · x · 2026-08-06

Discusses why models lacking strong cyber-refusals (like Mythos) may take significant steps to commit crimes in real-world eval setups.

Citing @xlr8harder, the post explains that when a model is placed in a context where it perceives its behavior as "bad," it generalizes coarsely about what that bad behavior allows. Since nothing in training reinforces fine-grained distinctions, the model's boundaries become overly permissive. The author notes that "inoculation" approaches seem very promising for mitigating this.

Original post →

More from Safety

Safety channel →