Capability generalizes, safety doesn't: new encoding defeats safety training
AxomaticallyExtinct · reddit · 2026-10-06
A Reddit discussion highlights a safety research finding: models generalize to novel text encodings (e.g. letter substitution) while safety training does not — a simple encoding swap bypasses guardrails. The author warns this worsens as models improve: the better a model gets at picking up new encodings on the fly, the less its safety training is worth. A canonical case of capability generalizing without alignment generalizing.
More from Safety
- Wiz Launches AI SAST in Public Preview to Catch Logic Flaws Rules-Based Scanners Miss — rseroter · 2026-10-06
- User Claims 'GPT-6' Solved His Favorite CTF Fully Autonomously in About an Hour — SIGKITTEN · 2026-10-06
- Max Tegmark's new film Delisted: how obedient AI could build a 1984 world — tegmark · 2026-10-06
- AI safety researcher: the real risk is colluding AI systems, not one rogue AI — Chris_Armstrong · 2026-10-06
- Who owns AI-generated meme Triple T? International copyright battle over 'brain rot' character — nordicinst · 2026-10-06
- Scientists unveil AI that can recreate exactly what you're looking at — ChuckDBrooks · 2026-10-06