AgentArk, distilling multi-agent debate into a single LLM, accepted to NeurIPS 2026
jindong_wang92 · x · 2026-09-25
AgentArk distills multi-agent debate dynamics into a single LLM's weights, converting test-time interactions into implicit capabilities. Key numbers: +4.8% avg accuracy over single-agent baseline, 120 experiments across Qwen3/Gemma 3/Llama 3, three distillation strategies (R-SFT, trajectory augmentation, process-aware distillation), 342K questions and 2M reasoning trajectories. Code is open source.
More from coding & agent
- pnpm urges devs to spend spare tokens fixing its 629 open issues, ships a ready-made agent prompt — itsOmSarraf_ · 2026-09-25
- Geoffrey Huntley Demos a Software Factory Where the Product Is Its Own IDE — teropa · 2026-09-25
- Open-Source iCloud MCP Runs Without a Mac: 31 Tools, Headless Chromium — sjdonado · 2026-09-25
- Microsoft ships enterprise product on OpenClaw after months of joint hardening work — steipete · 2026-09-25
- Open-sourced skill turns your codebase into a polished product promo video via Claude Code — op7418 · 2026-09-25
- BlackRock paper sparks debate: agent payments settle fine, but revoked permissions can't catch up — tallmetommy · 2026-09-25