Developer finds flaws in CritPt benchmark scores

scaling01 · x · 2026-09-02

A developer points out flaws in the CritPt benchmark, noting that the scores don't make much sense and fail to increase significantly with additional reasoning, suggesting the benchmark may not accurately reflect model capabilities.

Original post →

More from Models

Models channel →