Nonobench: 49 LLMs tested on nonogram puzzles, solve rate falls to 20% at 15x15

mauricekleine · reddit · 2026-10-04

Nonobench is an open benchmark testing how well LLMs solve nonograms (picross): each model gets the row and column clues once and must return the full grid — no tools, one attempt per puzzle.

Method: Standard mode has 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0); Hard mode uses ten random 20x20s verified to have unique solutions, five of which can't be solved by line logic alone. 130 variants across reasoning effort levels were run through OpenRouter or the labs' own endpoints.

Results: At each model's best effort level, solve rates drop from 85% (5x5) to 46% (10x10) to 20% (15x15). GPT-6 Astra solves all 30 Standard puzzles; on Hard mode, Claude Opus 5.5 solves 8 of 10 while 11 of 15 models solve none. Notably, when answers had to be one 400-character string, most models lost count before the logic got hard — hence Hard mode accepts an array of 20 row strings instead.

Limitations: one attempt per puzzle makes single results noisy (95% intervals shown). Site: nonobench.com; code is MIT-licensed on GitHub.

Original post →

More from Research

Research channel →