Tsinghua's AAArena finds AI agents hit rank 1 with replays but stall on complex games
rohanpaul_ai · x · 2026-10-11
Tsinghua researchers built AAArena, a benchmark from 12 games in their yearly bot-building contest, pitting a weight-frozen coding agent against 1,920 archived human programs. The agent reads rules, picks opponents, studies replays, and rewrites its bot within a match budget.
Key findings:
- Detailed replays beat win/loss-only feedback in all 3 tested games; a Pacman bot reached rank 1 with replays vs rank 11 without.
- More compute doesn't help: tripling the match budget failed to push any of 4 stuck bots to rank 1.
- Learning winning strategies from limited matches remains hard, especially against changing rivals.
Paper: arxiv.org/abs/2610.12341
More from Research
- Meta & CMU's IdeaScientist uses RL agents for cross-domain research ideation, lifting novelty from 36.3% to 67.0% — ZeYanjie · 2026-10-11
- Qwen, Kimi and GLM dropped full attention — 8 attention designs explained — julsimon · 2026-10-11
- OpenMed 3.0: Apache-2.0 clinical AI that runs fully local and never falls back to the cloud — dark-night-rises · 2026-10-11
- Google's TabFM: a zero-shot foundation model that predicts table data in one forward pass — Prompt Engineering · 2026-10-11
- 'Nothing new': researcher says engram overfitting is plain overfitting tied to over-parametrization — teortaxesTex · 2026-10-11
- dair-ai's Top AI Papers of the Week: CMU's Harness Learning and Six More Picks — dair_ai · 2026-10-11