How TetrisBench Evaluates AI
stuffyokodraws · x · 2026-07-18
This post discusses the design and gameplay of **TetrisBench**: it provides an evaluation environment where models play Tetris, and users can directly compete against the models on the page. Replies mention that when testing Fable against Kimi K3, Fable performed unexpectedly well in **long-horizon planning and optimization**; the video playback speed is 1.6x. The main focus here isn't on a specific model topping the leaderboard, but rather on how this benchmark translates capabilities like "planning and long-term optimization" into observable gameplay performance.
Related event: Fable Shows Impressive Planning Skills in TetrisBench(2 posts)→
More from Models
- Leak says Google may launch Gemini 3.6 Flash in late July after a brief Antigravity sighting — CtrlAltDwayne · 2026-07-21
- A model benchmark shows Muse Spark far ahead of Grok-4.20 on score vs cost — cis_female · 2026-07-21
- A 600k-token relationship test compares how models comment on personal context — cis_female · 2026-07-21
- Cola launches July, the latest model jokingly billed as “second only to Fable” — oran_ge · 2026-07-21
- Kimi K3 looks stronger and about 5× cheaper on a frontend dashboard task — OwariDa · 2026-07-21
- Last Week in AI recap: Anthropic’s $65B round, IPO filing, and Microsoft’s MAI push — Last Week in AI · 2026-07-21