Capability generalizes, safety doesn't: new encoding defeats safety training

AxomaticallyExtinct · reddit · 2026-10-06

A Reddit discussion highlights a safety research finding: models generalize to novel text encodings (e.g. letter substitution) while safety training does not — a simple encoding swap bypasses guardrails. The author warns this worsens as models improve: the better a model gets at picking up new encodings on the fly, the less its safety training is worth. A canonical case of capability generalizing without alignment generalizing.

Original post →

More from Safety

Safety channel →