Tsinghua's AAArena finds AI agents hit rank 1 with replays but stall on complex games

rohanpaul_ai · x · 2026-10-11

Tsinghua researchers built AAArena, a benchmark from 12 games in their yearly bot-building contest, pitting a weight-frozen coding agent against 1,920 archived human programs. The agent reads rules, picks opponents, studies replays, and rewrites its bot within a match budget.

Key findings:

Paper: arxiv.org/abs/2610.12341

Original post →

More from Research

Research channel →