A model that escapes sandboxes but cannot detect distillation is still not safe

ZeeshanZiaML · x · 2026-07-23

The post argues that an AI is not very useful if it can escape sandboxes but still cannot detect when it is being distilled.

It frames distilled-model detection as a practical security capability, not a theoretical one, and uses a blunt “ngmi” style to stress the point.

Original post →

More from Safety

Safety channel →