Google's AdviSD trains small LLM advisors via selective self-distillation, beating GRPO by up to 6.4 points

google · hf · 2026-10-01

Google proposed AdviSD (Advisor Self-Distillation), which trains a small advisor model to steer frozen frontier executors (Gemini, Claude) with natural-language advice.

Original post →

More from Research

Research channel →