325 AI Models Tried Building a Budget-Capped Monster Fighter; a Third Failed
Binxtv · reddit · 2026-10-09
For the Halloween project AI Monster Brawl, the team sent one identical prompt to all 325 text models on OpenRouter: design a fighter from a fixed parts menu within a 1000-point budget, with up to 12 if/then rules as its combat brain, output as a single JSON blueprint.
Results:
- 213 models (66%) passed on the first try, 88 qualified after retries, and 24 (7%) never produced a valid fighter.
- Budget overrun was the recurring failure: 91 of 136 rejected attempts overspent; 38 invented non-existent parts or conditions, 14 broke format rules, 7 returned unparseable JSON.
- On retries, models mostly kept the same combat rules and changed the build instead — about four in five retried qualifiers kept their rules unchanged.
The authors note it's a one-prompt snapshot rather than a benchmark and ask whether models should get a budget-checking tool.
More from Models
- FrontierCode Is a Private Cognition-Run Eval, Mistral Exec Clarifies — b_roziere · 2026-10-09
- Google ships a decision-making AI model into Chrome, tested against Gemini Nano and Decisions API — gaganghotra_ · 2026-10-09
- User feeds Grok Bot 60 seconds of screen recording, gets a surprisingly decent tutorial video — elonmusk · 2026-10-09
- Commenters slam Anthropic's vague usage policy: "the only meaning is they can ban you anytime" — ivan_bezdomny · 2026-10-09
- User slams codex as 'GPT-3.5-level': agent ignores updated analyses in knowledge base — Yamapama · 2026-10-09
- Confidence calibration beats raw accuracy: Nimble 9B automates 2.7x more routing traffic than Qwen 3.5-9B — AgitatedUsual8995 · 2026-10-09