MAPD distills agentic search into structured protocols and lifts Qwen3 scores
Junlin Liu · hf · 2026-07-28
- The paper tackles distillation for agentic search, where outcome-based RL gives sparse supervision and naive imitation leaks style rather than reasoning skill.
- It proposes Multi-Agent Protocol Distillation (MAPD): an offline multi-agent system decomposes queries, retrieves evidence, repairs failed searches, and converts traces into a structured JSON protocol.
- During training, this protocol is exposed only to a privileged branch of the student policy, creating a denser distillation signal alongside sparse RL.
- Across seven QA benchmarks, MAPD reports average success rates of 39.4% on Qwen3-1.7B and 44.4% on Qwen3-4B, while reducing style drift and verbosity degeneration.
More from coding & agent
- Claude Opus 5 helps a browser racing game run 4,300 sims at once — AaronMatthews25 · 2026-07-28
- JarvisHub turns a canvas into shared memory for multimodal creative agents — Yunlong Lin · 2026-07-28
- Codex and Claude Code browsers are being used to run mini apps as live agent tools — RileyRalmuto · 2026-07-28
- Practitioner says multi-agent workflows work better with 1–2 active agents, then long autonomous runs — edgarpavlovsky · 2026-07-28
- Only one of 63 API companies exposes all three agent-ready surfaces — ExtensionPea834 · 2026-07-28
- Reddit asks whether 136.7 million x402 settlements actually prove agent adoption — heiba_wk · 2026-07-28