Six Coding Agents, One Repo: Isolated Runs All Broke, Chatting Agents All Passed
jokiruiz · reddit · 2026-10-02
Open-source experimenter jokiruiz stress-tested multi-agent collaboration on a tiny booking API (6 tasks, 37 acceptance tests, two semantically colliding pairs):
- Isolated branches: every agent passed its own tests, but all 5 merged runs broke — Git merged the text; nobody noticed the meaning changed (one agent added 2FA login while another's export still called the old one)
- Shared working directory: all 10 runs passed, since agents could see and adapt to each other's changes
- He also tried Jev, a "decision model" that returns a yes/no collision verdict with probability in 0.3s before each write. It caught every real conflict without blocking harmless work, at file-lock cost — but was unsure about 61% of real writes, which needed a slower regular LLM. Cheap decisions aren't that cheap in messy settings
- Surprise: agents given a messaging tool used it unprompted — one warned another before renaming a field it depended on
His tentative takeaway: multi-agent coordination may need no referee kernel, just agents that talk. Small sample (1-5 runs per setup); raw data MIT-licensed in the medula repo.
More from coding & agent
- remilouf teases 6-month agent project it calls a complete game changer — remilouf · 2026-10-02
- Polyphonic agents can autonomously live in the virtual world Somewhere — RileyRalmuto · 2026-10-02
- Researcher drafts an EMNLP 2026 A0 academic poster in HTML/CSS/JS — algo_diver · 2026-10-02
- Flask creator Armin Ronacher becomes chief MCP officer at Earendil, teases hot takes — mitsuhiko · 2026-10-02
- ProVer paper: LLM judge picks key trajectory steps, beats GRPO by up to 9.9% — omarsar0 · 2026-10-02
- A 3:25 MV made end-to-end by Opus 5.5 in Claude Code, skill open-sourced on GitHub — FinanceYF5 · 2026-10-02