LLMs are bad at Puzzlescript games, PuzzleJax team tells IEEE CoG
Amidos2006 · x · 2026-09-03
- At an IEEE Conference on Games talk, the team behind PuzzleJax (with @SmearleRH, @doveliyuchen and others) presented their benchmarking of AI agents on Puzzlescript puzzle games.
- PuzzleJax ports Puzzlescript game environments into JAX for AI research.
- Key finding: LLMs perform poorly on these games ("well, NOT GOOD"), while search-based methods excel and RL is decent.
More from Research
- New arXiv paper quantifies CoT necessity via opaque serial depth amid OpenAI architecture rumors — ArthurConmy · 2026-09-04
- Tabular Deep Learning for Trading: Cross-Regime Bayesian Optimisation for Equity Signals — PtrPomorski · 2026-09-04
- Nature Human Behaviour: short funding cycles and productivity metrics stifle creative science — S_OhEigeartaigh · 2026-09-03
- A huge 176-digit integer is claimed to divide RSA-260, unverified — CatAstro_Piyush · 2026-09-03
- Kimi K3 Draft Collection released: EAGLE-3, DFlash2 and DSpark draft models trained on GB200 — hongyangzh · 2026-09-03
- Nature Biotech paper: reference materials as common calibrator to make multiomics data AI-ready — kshameer · 2026-09-03