Paper unveils preliminary Agent Swarm RL recipe; V4.1 teams hit 30% on ProgramBench
inductionheads · x · 2026-09-10
teortaxesTex hails a new paper that includes a preliminary Agent Swarm RL training recipe. He argues the official 20.3% ProgramBench score is far from the ceiling: a V4.1 team setup reaches 30%, well above the single-agent Sol and nearly 2x Kimi K3. He also speculates that OpenAI's interactive swarms may already exploit similar ideas.
More from coding & agent
- DSH V4.1 can now intelligently designate subagents, janky but fun — teortaxesTex · 2026-09-10
- Shrinking an agent's macOS tool surface from 84 tools to 19 cut schema context by 73% — bulutarkan · 2026-09-10
- How to run a hackathon entirely with ChatGPT Sites, all from your phone — gabrielchua · 2026-09-10
- KIE Nodes Next turns KIE.ai's API catalog into a native ComfyUI model library — felipederosilva · 2026-09-10
- 34-day agent run processed 21.5B input tokens for $200 with 98% cache hits — nodo48 · 2026-09-10
- Engineer stops reading AI-written code, shares 4 ways to keep judgment sharp — every · 2026-09-10