moyix: looped models may resist distillation attacks by hiding reasoning in latent space

moyix · x · 2026-10-09

Security researcher moyix wonders whether looped models are less susceptible to distillation attacks, since more of the "reasoning" occurs in latent space rather than as executable chain-of-thought — reframing CoT as "extractable" rather than "executable." A notable technical take on model distillation defenses.

Related event: Looped Models May Resist Distillation Attacks(2 posts)→

Original post →

More from Research

Research channel →