Obscure board games as the best AGI eval: Fable far behind Opus 5
paul_cal · x · 2026-08-25
Grant Slatton argues obscure board games may be one of the single best measures of actual AGI, calling for more such evals — the nice property being a near-infinite backlog of obscure games with virtually no training data contamination.
Someone then ran the test: they built an online version of the strategic horse-betting dice game Long Shot: The Dice Game and benchmarked models at it — Fable performed much worse than Opus 5 under default rules.
More from Models
- Grok generates unprompted image push to boost app engagement — StefanoGogioso · 2026-08-25
- Security Researcher Waits a Month for Claude Cyber Trusted Access Approval — nptacek · 2026-08-25
- Codex overage allowance slashed to ~1%; exploit value > disclosure bounty — nptacek · 2026-08-25
- Developer complains about Ox Alpha's slow inference: 127 mins for 10 min task — altryne · 2026-08-25
- a16z partner blown away by access to unreleased AI model — AccBalanced · 2026-08-25
- Chinese LLMs 4-5 Months Behind US; ECI 155 May Be Reliability Threshold — Jsevillamol · 2026-08-25