Using Internal Model Signals for Training Isn't Always Bad
gleech · x · 2026-07-19
The author discusses the classic "Most Forbidden Technique": whether using a model's internal representations as training signals is truly always undesirable.
The article argues that a blanket rejection of such methods is often exaggerated. While acknowledging that these practices are linked to "obfuscation" and require caution, the author contends that using internal model information shouldn't be equated directly with a bad practice—it depends on the specific conditions.
Reviewing relevant literature, particularly the contribution of The Obfuscation Atlas, the author plans to summarize the exact conditions under which using internal model signals during training is worth adopting.
More from Safety
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Follow-up: song name and year both optional in Spotify chatbot bypass — AaronBergman18 · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11
- Fields Medalist founds Mathematical AI Safety Institute to prove AI safe like cryptography — The Decoder · 2026-09-11
- DeepMind alignment researcher signs open letter urging coordinated AI slowdown — vkrakovna · 2026-09-11