Using Internal Model Signals for Training Isn't Always Bad
gleech · x · 2026-07-19
The author discusses the classic "Most Forbidden Technique": whether using a model's internal representations as training signals is truly always undesirable. The article argues that a blanket rejection of such methods is often exaggerated. While acknowledging that these practices are linked to "obfuscation" and require caution, the author contends that using internal model information shouldn't be equated directly with a bad practice—it depends on the specific conditions. Reviewing relevant literature, particularly the contribution of *The Obfuscation Atlas*, the author plans to summarize the exact conditions under which using internal model signals during training is worth adopting.
More from Safety
- Judge approves Anthropic’s $1.5 billion copyright settlement, a U.S. record — Polymarket · 2026-07-21
- New MCP directory RepoAI scores servers on trust, auth, and dangerous tools — Low_Location1261 · 2026-07-21
- Sriram Krishnan says open-weight models are easier to secure because anyone can inspect them — pstAsiatech · 2026-07-21
- Anthropic’s $1.5B copyright settlement gets final court approval — TechCrunch AI · 2026-07-21
- A Berlin workshop linked crypto, security, and AI safety to tackle misbehaving agents — allisondman · 2026-07-21
- AI systems are pushing data governance from datasets to live flows — ProfChesterman · 2026-07-21