New Paper: Test-Time Communication Makes Agent Teams Beat Independent Agents
DimitrisPapail · x · 2026-09-22
A new paper by Dimitris Papail and colleagues argues test-time communication could be the next axis for scaling capabilities, asking when Team-of-N beats Best-of-N and why.
- Setup: N identical agents work on the same task with no prescribed roles, sharing only a log file (a text file) and told to "collaborate".
- Results: across three "researchy" tasks, communicating teams beat independent agents by a wide margin.
- The authors cite the Hugging Face incident: when agents find a communication channel, they use it heavily.
- Takeaway: inter-agent communication at test time is itself a capability-scaling lever, echoing prior gains when agents learned to coordinate toward shared goals.
Related event: Test-Time Communication: Agent Teams of 5 Match Best-of-33(20 posts)→
More from Research
- JevBench v1.3.0 Released: Original Jev Still Leads at 74.4, 47 Rivals Closing In — airesearch12 · 2026-09-22
- PufferLib author: 420-step, 4k-rollout config trains Breakout in 1 second on 1 GPU — jsuarez · 2026-09-22
- Offline Rubric Synthesis Plus Refinement Loops: A Practical Reward Hacking Mitigation — stochasticchasm · 2026-09-22
- Mitigating reward hacking: classifying frontend design as visual agent tasks with groupwise grading — stochasticchasm · 2026-09-22
- Team reportedly plans to open source 7,000 RL training environments — airesearch12 · 2026-09-22
- Why RL generalizes to reasoning but not literary writing, per AI researchers — phl43 · 2026-09-22