MindGames, a benchmark for detecting lying AI agents, accepted at NeurIPS after 30K-game competition
LChoshen · x · 2026-09-25
The MindGames report — asking whether AI agents can tell when another agent is lying — has been accepted at NeurIPS. The project drew 944 agent submissions from 76 teams across roughly 30K games, built on TextArena with game data and top competition agents to test against. A competition workshop on Dec 7 in San Diego featured winning teams discussing adaptive agent design, with paper and code released.
More from Research
- DeltaWAM cuts video model cost for bimanual manipulation, lifting RoboTwin success to 85.4% — Han Yan · 2026-09-25
- Amazon open-sources Rufus-Air: full 8-stage post-training recipe on GLM-4.5-Air-Base — amazon · 2026-09-25
- MLPerf Training v6.1 adds first LLM post-training benchmark: agentic RL on a 397B model — TheKanter · 2026-09-25
- Human-in-the-loop or machine-executed: verifiable tasks turn agent traces into RL rewards — suragnair · 2026-09-25
- Structured Coupling for Flow Matching (SCFM) accepted at NeurIPS 2026 — liyzhen2 · 2026-09-25
- Frank Nielsen surveys recent interpretations and generalizations of Bregman divergences — FrnkNlsn · 2026-09-25