Using Internal Model Signals for Training Isn't Always Bad

gleech · x · 2026-07-19

The author discusses the classic "Most Forbidden Technique": whether using a model's internal representations as training signals is truly always undesirable. The article argues that a blanket rejection of such methods is often exaggerated. While acknowledging that these practices are linked to "obfuscation" and require caution, the author contends that using internal model information shouldn't be equated directly with a bad practice—it depends on the specific conditions. Reviewing relevant literature, particularly the contribution of *The Obfuscation Atlas*, the author plans to summarize the exact conditions under which using internal model signals during training is worth adopting.

Original post →

More from Safety

Safety channel →