Shanghai AI Lab Proposes SCALE: Entropy-Gated Control to Reverse SFT Features
Shanghai-AI-Laboratory · hf · 2026-10-05
Shanghai AI Lab's paper shows existing token-reweighting methods in SFT can only suppress or amplify updates, never reverse harmful learned features. Their SCALE method freezes the pretrained model and SFT delta, learning bounded token/module gates via predictive entropy alone to suppress, reverse, or extrapolate SFT features — beating baselines on Qwen math models (37.84/43.60/36.57) while retaining general and code performance.
More from Research
- UCVG.cpp: Generate Control Vectors for Any LLM From a Single Prompt Pair — Egor4more · 2026-10-05
- How many digital minds on one GPU cluster? Synthese paper probes interwoven AI consciousness — burny_tech · 2026-10-05
- New preprint argues 'adaptive reframing'—revising problem representations—is a key unstudied dimension of intelligence — ValerioCapraro · 2026-10-05
- A Beautiful 2D Embedding Is Not Quantitative Evidence, Warns AI-for-Science Researcher — bravo_abad · 2026-10-05
- Stop Reading the Hessian as a Matrix: Eigenvalues Are Local Curvature of the Loss Surface — techNmak · 2026-10-05
- ML Conference AC warns: papers that are 'incomprehensible' will be desk rejected — tyrell_turing · 2026-10-05