Google's AdviSD trains small LLM advisors via selective self-distillation, beating GRPO by up to 6.4 points
google · hf · 2026-10-01
Google proposed AdviSD (Advisor Self-Distillation), which trains a small advisor model to steer frozen frontier executors (Gemini, Claude) with natural-language advice.
- Key finding: corrections that plausibly look right but don't change execution can hurt learning; the paper proves this for shared-parameter advisors
- Method: outcome-based RL plus selective self-distillation from a feedback-conditioned advisor copy — the advisor scores the same executor response with and without its advice, supervising only where the score gap is large; no executor likelihoods or extra rollouts needed
- Qwen3-8B advisors outperform advisor-GRPO by 4.2-6.4 points on BFCL-v3 and 3.9-5.1 points on EnvScaler
- Advisors generalize out-of-domain and transfer across executor versions and model families
More from Research
- Pantheon open-sources Argus, a SoTA robotics data annotator that audited 9 datasets and released 3,546 episode annotations — silver__tsuki · 2026-10-01
- Blender Bench v1: GPT-6 Astra leads 3D benchmark, Claude Opus 5.5 wins modeling and cloth — yshan2u · 2026-10-01
- Cambridge's SAM-Brain3D surpasses existing methods for early brain disease detection — lawrennd · 2026-10-01
- Protein binding folding models face data starvation after PDB has been fully juiced — anshulkundaje · 2026-10-01
- Stanford's UniEvo-VL Self-Distillation Lifts Qwen-image GenEval From 0.747 to 0.808 — stanfordnlp · 2026-10-01
- NVIDIA's Mid-Harness Scales Actions at the Model-Harness Boundary, Lifting TerminalBench Pass@1 to 68.03% — nvidia · 2026-10-01