Category-aware expert RL framework hits 59% on SWE-bench Multilingual
Logics-MLLM · hf · 2026-09-22
A new paper tackles category see-saw in repository-level SWE agents, where pooled agentic RL improves some task categories while regressing others. Key points:
- An executable task construction pipeline and SWE Labeler (evidence-grounded multi-axis labeling) organize training pools.
- Same-origin category experts alternate long-horizon Agentic-miniRL with Refresh-Repair-Expand (RRE): refreshing instance mastery, reusing verified successful trajectories for Repair SFT, and reselecting tasks for further RL.
- Label-routed multi-teacher on-policy distillation (MOPD) consolidates experts into one deployable student, with ReLU-gated reward extrapolation keeping only each teacher's improving direction — no external model needed for trajectories.
- The final policy reaches 58.04% on Pro-618 and 59.00% on SWE-bench Multilingual, +5.39 and +2.78 points over the base model.
More from coding & agent
- OpenAI shows how a surgeon uses Codex to build tools and review PubMed literature — OpenAIDevs · 2026-09-22
- Similarity Isn't Relevance: Four Layers Every Personal Agent Memory System Must Separate — sujingshen · 2026-09-22
- JevHarness lets an LLM author and freeze task-specific agent harnesses — but who owns the frozen SOP? — sujingshen · 2026-09-22
- Self-modifying harnesses: who grades the exam matters more than the evolution — sujingshen · 2026-09-22
- 15 Claude Skills Worth Installing, Plus a Chat-Prompted Blender Animation Pipeline — Div_pradeep · 2026-09-22
- KeycardAI CEO: AGI isn't the bottleneck for agents — trust, identity and permissions are — brucemacv · 2026-09-22