Sherpa: Multi-Turn RL Makes LLM Teachers Boost Student Scores by 20.5 Points
SALT-NLP · hf · 2026-10-08
- SALT-NLP introduces Sherpa, a multi-turn reinforcement learning framework built on the insight that solving a problem differs from teaching it, and that effective teaching varies per learner.
- Method: instantiate multiple LLM-based student archetypes with distinct learning preferences, then train a teacher model to adapt instruction by directly maximizing simulated students' learning outcomes.
- Results: instructed students improve by an average of 20.5 percentage points across all archetypes; on MathTutorBench, overall pedagogy score rises from 52.5% to 79.2%; human studies prefer the trained teacher in 79.6% of pairwise comparisons.
- The authors frame this as a step toward AI tutors teaching real students.
More from Apps
- OpenAI DevDay community turns ModRetro Chromatic handheld into a Codex-powered game dev platform — OpenAIDevs · 2026-10-08
- Microsoft ships Windows hybrid intelligence: MXC agent sandbox GA, HydraFusion local-cloud models — pavandavuluri · 2026-10-08
- OpenAI rolls out ChatGPT for Teens with default protections, previews College Planner — OpenAI · 2026-10-08
- Delphi launches Predict Anything, a forecasting agent that gives odds on any claim — benfielding · 2026-10-08
- Kuliko launches an MCP study workspace letting ChatGPT and Claude manage flashcards and quizzes — Bright-Pound · 2026-10-08
- Grok Bot v0.68.1 Adds Collaborative Slide Deck Creation and Formatted Emails — mark_k · 2026-10-08