Tencent releases GameXpert-Bench to evaluate coding agents in game development
Tencent-Hunyuan · hf · 2026-08-25
Tencent released GameXpert-Bench to evaluate the capabilities of coding agents in professional game development. Covering three key stages—generation, repair, and optimization—the benchmark uses interactive and behavioral tests. It reveals that agents excel at building playable foundations but show weaknesses in defect discovery and regression preservation.
More from coding & agent
- Long-Horizon Agent Dev Pain: Not Enough Time to Run Full Rollout — agihouse_org · 2026-08-25
- AGI House Hackathon Recap: Long-horizon Agents Need Verification, Not Just Memory — agihouse_org · 2026-08-25
- Terminal-Bench 3.0: Top Model Score Plummets to 43.5% as Agents Fail to Fix Root Causes — ajratner · 2026-08-25
- Spine-Branch Framework Boosts Multi-Agent Success by 16.5% — ZhiyuChen4 · 2026-08-25
- Anthropic Showcases Creative Claude Code Projects: From Medical Viewers to 3D Streets — 机器之心 · 2026-08-25
- AI Agents Need Four Types of Memory to Mimic Human Capabilities — _jaydeepkarale · 2026-08-25