Detective game benchmark: agent swarm plus strong reasoner solves 96/100 cases
No_Yogurtcloset_7050 · reddit · 2026-10-11
The LMGame team built a procedurally generated Victorian London detective game (100 addresses, 10 suspects, 12 case-hour limit) to test whether small fast agent swarms can solve murder cases.
Results (standard / frame-up):
- GPT6-Astra Holmes + 12 Jev scouts: 96 / 93
- Rule-based Holmes + swarm: 95 / 0 (clue-counting fails when witnesses lie)
- Lone GPT6-Astra: 57 / 59
- Leaderless 100-agent voting swarm: 38 / 42, median 2.2 case-hours
Takeaways: swarms cover evidence, a strong reasoner arbitrates; leaderless swarms reach premature consensus; with honest evidence a rule-based leader matches GPT6-Astra cheaply, but only a model that questions evidence provenance survives frame-ups (93/100).
Open source with case replay and other environments: github.com/lmgame-org/Gaming-JevSwarm
Related event: Small-Model Agent Swarm Solves 96% of Detective Game Cases(2 posts)→
More from coding & agent
- Synthesia launches Syren Video: prompt-to-video agent powered by Opus 5.5, free to try — heyshrutimishra · 2026-10-11
- Codex lead jokes 'the day we reach perfection it will be resets from there onwards' — xiaohu · 2026-10-11
- One MCP server for X, Instagram, WhatsApp, Telegram and Gmail with per-action approvals — gauthi3r_XBorg · 2026-10-11
- Dev ports Sega MegaDrive game natively in under a day with single Codex/Claude session — ssh4net · 2026-10-11
- Grok Bot Negotiated His Internet Bill Down $15/Month Over Live Chat — jarrodwatts · 2026-10-11
- banteg surprised Codex cloud auto-configures its environment just by reading the repo — banteg · 2026-10-11