TALES Benchmark on LLM Game-Playing Accepted to NeurIPS 2026, New Results Coming
tw_killian · x · 2026-09-29
Researcher ccui9 announced that TALES, a benchmark asking how well LLMs can play games, has been accepted to NeurIPS 2026. The arXiv paper is slightly out of date: the team has since gathered results on a wave of new models, including Astra and Opus 5.5, and plans to update the paper soon.
More from Research
- Researchers flag LLM checking limits: validation is fluent, not formally verified — anshulkundaje · 2026-09-29
- Foresight hosts SF conference on AI-first science with DeepMind, MIT speakers — juanbenet · 2026-09-29
- MetaLint: Qwen3-4B lifts code lint detection F-score 2.7x to 70.4%, matching o3-mini — dan_fried · 2026-09-29
- How do you make a chess bot blunder believably? Mixing Maia and Stockfish isn't enough — space64-llc · 2026-09-29
- Aether AI releases 16B CausalWM, a world model that reasons about causality before generating future frames — jiqizhixin · 2026-09-29
- LT-OPD On-Policy Self-Distillation Lifts 5% Visual-Token Retention to 82.3% — Junxian Li · 2026-09-29