Open benchmarks are broken: models regex out offloaded answers, researcher argues

a1zhang · x · 2026-09-23

a1zhang argues open benchmarks have become frustrating because it's hard to tell what models actually know. On 'new' long-context benchmarks, some RLMs simply regex for offloaded information and grab the solution.

Original post →

More from Models

Models channel →