Closed-form weight surgery transfers 4B capabilities into 0.8B with 4 anchor blocks

AdventurousTwo6445 · reddit · 2026-10-06

Instead of weeks of token-heavy distillation, the author extracts layer-to-layer hidden state trajectories on a handful of calibration prompts and solves closed-form weight updates in the student's SwiGLU MLP blocks.

Key findings:

Code, benchmark logs and weights are open-sourced (GitHub DynamicTune; HF F-Labs/Qwen3.5-0.8B-DynamicTune-Base). Proposed next tests include transplanting reasoning from Qwen3.8-27B to smaller models and unknotting layers 1-22 with SAEs.

Related event: Closed-form weight surgery transplants 4B model capabilities into 0.8B(2 posts)→

Original post →

More from Research

Research channel →