Learn2Play Bench tests LLM agents learning novel game rules from scratch via interaction
NationalUniversityofSingapore · hf · 2026-10-09
- NUS introduces Learn2Play Bench: newly designed text games with novel or counterintuitive rules, so agents must learn through interaction rather than pretrained knowledge; reproducible feedback, automatic scoring, and unseen game instances test transfer.
- Three findings: (1) retaining complete raw records of actions and feedback supports better learning than summarizing experience into rules; (2) top human players beat evaluated agents with more varied strategies and less repetition; (3) with the backbone fixed, swapping the harness improves performance while cutting estimated inference cost.
More from coding & agent
- autoicd-mcp ships automated ICD-10 medical coding MCP server with 74,000+ code search — modelcontextprotocol · 2026-10-09
- pohjola-api: agent-native Finnish company data API priced at $0.01 per call via x402 — modelcontextprotocol · 2026-10-09
- Every model looked bad in my eval — the bug was my answer key, not the models — jgarg27 · 2026-10-09
- Pi has no official Subagents, but 4 community extensions emerged; author shares extension stack order — solyarisoftware · 2026-10-09
- Machine Desktop launches: every agent gets its own cloud computer and routines that run while you're away — tsi_org · 2026-10-09
- badclaude, the viral parody AI project, is now open source via npm — dotey · 2026-10-09