Developer finds flaws in CritPt benchmark scores
scaling01 · x · 2026-09-02
A developer points out flaws in the CritPt benchmark, noting that the scores don't make much sense and fail to increase significantly with additional reasoning, suggesting the benchmark may not accurately reflect model capabilities.
More from Models
- OpenAI previews Astra, a cybersecurity model reaching 'Critical' threshold — mikegiannulis · 2026-09-02
- Dev: Anthropic 5.1 is a massive leap over previous model — doodlestein · 2026-09-02
- Fable 5.1 test fails to meet expectations — Angaisb_ · 2026-09-02
- Claude Fable 5.1 generates a full Mario Kart game with a single prompt — ezshine · 2026-09-02
- Gemini 3.7 Flash Demonstrates Agentic Video Understanding — otarU · 2026-09-02
- OpenAI's Astra achieves 100% success rate on ExploitBench vulnerability tests — JiaweiLiu_ · 2026-09-02