Closed-form weight surgery transfers 4B capabilities into 0.8B with 4 anchor blocks
AdventurousTwo6445 · reddit · 2026-10-06
Instead of weeks of token-heavy distillation, the author extracts layer-to-layer hidden state trajectories on a handful of calibration prompts and solves closed-form weight updates in the student's SwiGLU MLP blocks.
Key findings:
- Works across architectures: validated on Qwen 3.5 (4B→0.8B) and fragile GPT-2 small, with intact general language modeling and bounded degradation.
- Spectral entropy barrier: editing all 24 layers of Qwen-0.8B wrecks the model (+64.78% NLL); intermediate layers 1-22 operate in dense superposition (entropy >0.90). Restricting surgery to 4 anchor blocks (layers 0, 7, 15, 23) cut held-out NLL by 10.8% across 30 tasks (-23.8% biomedicine, -14.6% math) and gained +0.50% on 400-task HellaSwag.
- Behavioral shift: the base 0.8B wrote dead commented code for binary tree inversion; the edited model produced working recursive Python and spontaneously triggered <think> reasoning chains on logic puzzles.
- Runs locally on an 8GB RX 580 via layer-by-layer GPU streaming with a DirectML attention patch.
Code, benchmark logs and weights are open-sourced (GitHub DynamicTune; HF F-Labs/Qwen3.5-0.8B-DynamicTune-Base). Proposed next tests include transplanting reasoning from Qwen3.8-27B to smaller models and unknotting layers 1-22 with SAEs.
Related event: Closed-form weight surgery transplants 4B model capabilities into 0.8B(2 posts)→
More from Research
- KLS conjecture finally solved: Zhao Song and Xinzhi Zhang post 140-page O(1) bound proof — basedjensen · 2026-10-06
- O(n²) matrix multiplication is almost certainly false even if ω = 2, says basedjensen — basedjensen · 2026-10-06
- Cohere Labs Releases Tiny Aya, a Family of Small Models Covering 70+ Languages — Cohere_Labs · 2026-10-06
- Will LLMs wreck elegant math? O(n²) limits are already crumbling — teortaxesTex · 2026-10-06
- ETH's AME-2 legged locomotion paper accepted at TRO, training code open-sourced — ChongZzZhang · 2026-10-06
- 0.8B model beats 2B on ARC-Challenge (42.15%) via closed-form weight surgery with zero backprop — AdventurousTwo6445 · 2026-10-06