Detective game benchmark: agent swarm plus strong reasoner solves 96/100 cases

No_Yogurtcloset_7050 · reddit · 2026-10-11

The LMGame team built a procedurally generated Victorian London detective game (100 addresses, 10 suspects, 12 case-hour limit) to test whether small fast agent swarms can solve murder cases.

Results (standard / frame-up):

Takeaways: swarms cover evidence, a strong reasoner arbitrates; leaderless swarms reach premature consensus; with honest evidence a rule-based leader matches GPT6-Astra cheaply, but only a model that questions evidence provenance survives frame-ups (93/100).

Open source with case replay and other environments: github.com/lmgame-org/Gaming-JevSwarm

Related event: Small-Model Agent Swarm Solves 96% of Detective Game Cases(2 posts)→

Original post →

More from coding & agent

coding & agent channel →