GPT-6 Astra is first model to solve all 30 puzzles in open-source nonogram benchmark
mauricekleine · reddit · 2026-09-29
A Reddit user has benchmarked LLMs on nonograms since January: one attempt per puzzle, no tools, answers verified against row/column clues. GPT-6 Astra (xhigh) is the first model to solve all 30 Standard puzzles (5x5–15x15); back in January the best model solved only 3 of 10 15x15s.
- On the new Hard mode (ten random 20x20s), Astra solves 5/10 — all five solvable line-by-line, none requiring deeper search
- Claude Opus 5.5 leads Hard mode with 8/10; 11 of 14 models couldn't solve any 20x20
- Everything is open source at nonobench.com
More from Models
- Dev after 2 days: Claude is excellent, Codex great for long-horizon tasks but poorly designed — cneuralnetwork · 2026-09-29
- Opus 5.5 tops Drone-Bench and cheats far less than prior Claude models — scaling01 · 2026-09-29
- User claims 'Opus 5.5' turned a post on agent harnesses into an explainer video in one shot — alex_verem · 2026-09-29
- Engineer proud as Sonnet 5.5 scores 61.6% on chartography benchmark — echen · 2026-09-29
- ProgramBench multi-agent eval: Opus 5.5 fastest with a 5-agent team, Sonnet 5.5 with subagents — jyangballin · 2026-09-29
- Arrow 2 Telos tops Design Arena's SVG generation benchmark — AWizardWhoCodes · 2026-09-29