Test-time communication emerges as next scaling axis as compute shifts from pretraining
DimitrisPapail · x · 2026-09-21
Sarah Hooker amplifies a new paper by Dimitris Papail et al. arguing compute is shifting from pretraining—where FLOPs yield diminishing returns—to test-time compute, requiring fundamentally different infrastructure; the first wave of agent-designed infra is arriving.
The paper asks when communicating agent teams beat independent agents (Team-of-N vs Best-of-N). N identical agents with no prescribed roles, collaborating only via a shared text log, substantially outperformed independent agents across three research-style tasks. The authors position test-time communication as a next axis for scaling capabilities.
More from coding & agent
- MechFaber: Claude Code designs a 99-part quadruped with firmware co-simulated in Renode and MuJoCo — SpeedyBrowser45 · 2026-09-22
- Exa MCP hits 5,000 GitHub stars as AI agents flock to its search integration — TheIshanGoswami · 2026-09-22
- 670,000 agent skills, no trust layer: bot scan finds 69% never reliably fire — markjeffrey · 2026-09-22
- Eight disruptive use cases for Jev: from millisecond evals to AI guardrails — nkmrao · 2026-09-22
- Training on production traces: single-trajectory RL may unlock continual learning — rhythmrg · 2026-09-22
- Anthropic's Swiss cheese model explains why passing evals isn't enough for agents — hugobowne · 2026-09-22