Nonobench: 49 LLMs tested on nonogram puzzles, solve rate falls to 20% at 15x15
mauricekleine · reddit · 2026-10-04
Nonobench is an open benchmark testing how well LLMs solve nonograms (picross): each model gets the row and column clues once and must return the full grid — no tools, one attempt per puzzle.
Method: Standard mode has 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0); Hard mode uses ten random 20x20s verified to have unique solutions, five of which can't be solved by line logic alone. 130 variants across reasoning effort levels were run through OpenRouter or the labs' own endpoints.
Results: At each model's best effort level, solve rates drop from 85% (5x5) to 46% (10x10) to 20% (15x15). GPT-6 Astra solves all 30 Standard puzzles; on Hard mode, Claude Opus 5.5 solves 8 of 10 while 11 of 15 models solve none. Notably, when answers had to be one 400-character string, most models lost count before the logic got hard — hence Hard mode accepts an array of 20 row strings instead.
Limitations: one attempt per puzzle makes single results noisy (95% intervals shown). Site: nonobench.com; code is MIT-licensed on GitHub.
More from Research
- NYU's $5 open-source eFlesh magnetic tactile sensor is 3D-printable — lukas_m_ziegler · 2026-10-04
- NeurIPS paper finds vision models fail to recognize objects without their typical neighbours — FrancescoLocat8 · 2026-10-04
- SPEAR Simulator Opens Up 14K+ Unreal Engine Functions for Embodied AI Research — rsasaki0109 · 2026-10-04
- Chalmers builds closed-loop AI scientist that autonomously makes experimentally validated biology discoveries — Dr_Singularity · 2026-10-04
- COLM 2026 to host LSEI workshop on Oct 9 exploring how LMs learn by interacting with world and agents — berkeley_ai · 2026-10-04
- DNA Rewrites 400-Year-Old Maya Legend: All 64 Sacrificed Children at Chichén Itzá Were Boys — aakashgupta · 2026-10-04