TAOCP open problems released as a dataset to benchmark frontier models
sytelus · x · 2026-09-05
- Author released taocpopenproblems on Hugging Face, a dataset of open problems from all volumes of Knuth's The Art of Computer Programming (<1K entries, parquet, math/algorithm tags).
- He argues the only interesting benchmark for frontier models going forward is solving open problems, making high-taste problem curation crucial.
- The dataset tracks per-problem status: some solved and independently reproduced (with verification scripts), others still open, with citations to recent literature.
More from Models
- Reviewer: OpenAI's GPT-6-Astra finally 'gets what you mean,' with Fable-level intelligence and real gains in game dev — pvncher · 2026-09-05
- Meta ships Muse Spark 1.3 with max reasoning, pitching frontier performance at non-frontier prices — AIatMeta · 2026-09-05
- LLMs are now making up words that don't exist, not just jargon — StewartalsopIII · 2026-09-05
- Alexandr Wang teases Muse Spark 1.3 with a max reasoning level after safety testing — alexandr_wang · 2026-09-05
- Muse Spark 1.3 max publicly released with stronger coding and agentic performance — jordihays · 2026-09-05
- Muse Spark 1.3 max released with significantly stronger coding and agentic performance — alexandr_wang · 2026-09-05