GPT-6 Astra posts first perfect 30/30 on Nonobench; open models score 0/10 on 20×20 Hard mode
mauricekleine · reddit · 2026-09-27
Nonobench v1.2 tests 43 LLMs on nonogram logic puzzles, with all prompts and outputs public.
- GPT-6 Astra scores 30/30 on 15×15 puzzles — the first perfect run on the benchmark.
- Best open weights: DeepSeek V4 Pro at 83% (tied 4th); DeepSeek V4.1 Flash hits 77% for $0.84 total.
- New 20×20 Hard mode (10 unique-solution puzzles, answered row by row after a community finding that most models miscount 400-character strings): Opus 5.5 takes 8/10; every open-weight model scores 0/10.
OpenRouter-only for now; code, raw data, and API on GitHub.
More from Models
- Musk confirms Grok 'upgrades' as users notice dramatic speed boost — elonmusk · 2026-09-28
- "System 2 models built the brain, but System 1 is building the nervous system" — ai · 2026-09-28
- TeleOCR Trends on Hugging Face: A Qwen2.5-VL-Based Chinese Document OCR Model — XingChen-AGI · 2026-09-28
- Kaggle Game Arena: Evaluating LLMs via Head-to-Head Chess, Poker, and Werewolf — kaggle · 2026-09-28
- Perplexity CEO: still using sol 6 for knowledge work — cheap, fast, great compaction — gabriel1 · 2026-09-28
- NerfBench's First Results Find No Nerf: Claude Opus 5.5 Dips Just 0.8% vs Launch — alejandroll10 · 2026-09-28