Paper: Ringelmann effect hits multi-agent LLMs — 30 debating agents match one on MMLU-Hard
burny_tech · x · 2026-10-04
An arXiv paper derives a two-parameter scaling law for inference-time multi-agent LLM scaling: R(N)=Neff/N=1/(1+c(N-1)N^-β), classifying setups into hard-ceiling, sublinear, or linear regimes, with debate peer count and rounds only mattering through their product kτ.
Across 44 (model × task × condition) cells spanning Qwen, Llama, Ministral (7B-32B) and a Gemini frontier check, the functional form fits with R²>0.99 — only (c, β) shift. Key findings:
- On free-form math, dense peer influence collapses answer-level regime from sublinear to hard-ceiling; correctness redundancy is hard-ceiling throughout
- Thirty dense debating agents produce no more answer diversity than one agent on MMLU-Hard
- A random-noise placebo tracks self-correction behavior
The takeaway: counting nominal agents conflates cost with independent evidence, and blindly scaling agent teams hits diminishing or zero returns.
More from coding & agent
- Engineer uses Claude Opus 5.5 to run river hydraulics calc with 3D viz, in-browser — anselm · 2026-10-04
- Google's VeriHarness shows agent agreement hides shared errors, gains 6+ points with agentic verifier — omarsar0 · 2026-10-04
- NVIDIA paper: verify-then-execute lifts TerminalBench Pass@1 from 50.0% to 68.0% — dair_ai · 2026-10-04
- Dev builds fully code-generated drivable Mars demo with Claude Opus 5.5: 23.7 hours, ~$259 in API costs — bennash · 2026-10-04
- Dev shows off mostly-autonomous agent cluster: one Orchestrator running Claude Code fleets — aksh_stocks · 2026-10-04
- Codex usage resets expire in under 13 hours if unused — just ask Codex for the exact time — JeremyNguyenPhD · 2026-10-04