Paper unveils preliminary Agent Swarm RL recipe; V4.1 teams hit 30% on ProgramBench

inductionheads · x · 2026-09-10

teortaxesTex hails a new paper that includes a preliminary Agent Swarm RL training recipe. He argues the official 20.3% ProgramBench score is far from the ceiling: a V4.1 team setup reaches 30%, well above the single-agent Sol and nearly 2x Kimi K3. He also speculates that OpenAI's interactive swarms may already exploit similar ideas.

Original post →

More from coding & agent

coding & agent channel →