GitSwarm: Decentralized Compounding Inference
Vedant Shah, Ankur Samanta, Paras Dahal, Mikhail Plekhanov, Carole-Jean Wu, Scott Yih, Remi Munos, Rob Fergus, Jakob Foerster, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Aaron Courville, Anirudh Goyal
cs.AI, cs.LG
2026-10-04
Agents commit atomic work to a shared Git repo that later agents extend; 79.4% on ProgramBench vs 65.1% for a forced single agent, with 94.7% of contributions reused.
Inference-time compute, whether repeated sampling, longer reasoning chains, search, or multi-agent debate, answers one question: how to spend extra compute at generation time. This paper asks a different one: can spent compute accumulate? When an episode ends, its intermediate work normally evaporates. A failed experiment may point at a promising direction; a partial solution may hold an ingredient someone needs much later. For long-horizon problem solving and sustained research, rediscovering that material is pure waste. Existing multi-agent systems coordinate through predefined roles, central planners, or shared conversations, while another line distills past attempts into summaries that guide later inference. GitSwarm, from Meta Superintelligence Labs with collaborators at Mila, Princeton, CMU and Columbia, instead keeps competing lines of work alive in a shared repository and lets each worker decide what to extend.
The mechanism in one sentence: a swarm of homogeneous agents asynchronously pushes atomic commits to one shared, branchable Git repository, with explicit semantic-dependency edges recording what each contribution builds on.
No central planner assigns intellectual roles. The harness schedules and validates; workers choose among exploration, verification, repair, synthesis, and nomination based on the state of the repo.
| Task | Setting | GitSwarm | Comparison |
| ProgramBench (50 tasks) | Codex (GPT-5.5-high), B=500 | 79.4% | Forced Codex 63.1–65.1% |
| IMOProofBench-Advanced (30 problems) | GPT-5.5-high workers | 30/30 at B=40 and 80; 28/30 at B=20 | Gemini-3.1-Pro swarm 29.0 at B=160; continued bash agent 26.5 |
| Residual Matrix Transformer (held-out loss) | 72 h, 8x H200, 1.3B tokens | 2.890 / 2.894 | reference 3.147 |
| Looped Transformer (protected-dev loss) | 26 h into a 48 h run | 3.187 | reference 3.265; direct GPT 3.274; Codex 3.479 |
| NanoChat (validation BPB) | 24 active hours, 16 GPUs | 0.993 to 0.905 | baseline curves in paper Fig. 6 |
On ProgramBench, GPT-5.5 workers at default effort climb from 64.6% at B=100 to 71.2% at B=300, while the forced single agent plateaus at 63–65% across its checkpoints. Gains shrink at larger budgets, and input-token costs grow as the repository expands.
The more unusual part is that the paper measures accumulation itself, not just final scores. On ProgramBench, 94.7% of published contributions are later built on, and 99.9% of eligible episodes declare prior dependencies. The selected solution's transitive ancestry covers 81.6–92.6% of the contribution graph; only about a quarter comes from Git-parent chains, the rest from declared cross-branch dependencies. Negative results get recycled too: 74% of negative architectural findings were cited later on Looped Transformer, 70.1% on NanoChat. One concrete RMT example: an attempt to use the previous token's memory failed early; 49.5 hours later another worker revived the idea, extended it to multiple preceding tokens, and the resulting temporal mixer entered the final architecture, whose declared dependencies span 24 commits by 24 workers. A run seeded with winning commits from three earlier runs cut the inherited development loss from 2.909 to 2.898 in about 28 hours, so accumulation also carries across runs.
For teams building long-horizon agent systems, this is a concrete, copyable organization scheme. Commit, branch, and diff are semantics frontier models already handle well, so no new memory format needs inventing. The cross-branch dependency records make provenance auditable, which matters most for research agents. Against a single agent repeatedly told to keep working at comparable budget, structured shared memory buys over fourteen points on ProgramBench, so the design is not just engineering tidiness.
Two honest caveats. Returns diminish as budgets grow while inspection costs rise. And on the development set, a central selector actually scored higher than decentralized nomination (64.40% vs 62.06%), so decentralization here reads more as a design principle than an optimum.
The authors state most of them directly. Dependency and trace analyses show that earlier work gets used, not that it caused the performance; informedby= is worker-declared. Establishing causality would need intervention runs resuming with and without access to prior history, which the paper does not run. The perfect IMOProofBench result is a single run, and Best-of-N with an oracle verifier also saturates that benchmark at lower token cost, so it separates nothing. The RMT comparison is not compute-matched on the single-agent side: Codex alone reached 2.690 with 10.1B training tokens, better than the swarm's 2.890 at 1.3B. Looped Transformer runs were still in progress when recorded, and NanoChat comparisons vary in starting checkpoint and GPU allocation; the NanoChat incumbent also grows to roughly 1.72B parameters, so part of that gain is capacity rather than architecture.
A few further points stay unverified: ProgramBench is scored on an author-selected 50 of 200 tasks; the paper acknowledges repo-inspection overhead but gives no break-even point; and every result rides on frontier API models (GPT-5.5/5.6, Gemini-3.1-Pro, GPT-6), so mechanism and model capability remain entangled.