Critic says CritPt and SciCode evals are broken, results untrustworthy until fixed
inductionheads · x · 2026-09-04
In a debate with TheZvi, teortaxesTex argues that both CritPt and SciCode benchmarks are broken, that there is genuine weakness in knowledge (and perhaps GDP-related tasks) versus Anthropic, and concludes results are untrustworthy until the evals are fixed.
More from Models
- Codex Down for Much of Rollout Day — and Users Not Getting New Model Either — RexDouglass · 2026-09-04
- Neuralese explained: why OpenAI's Astra architecture has safety researchers alarmed — ShakeelHashim · 2026-09-04
- K2 Horizon launches six fully open models from 0.9B to 375B with training code, data recipes and logs — aliscodes · 2026-09-04
- Is Astra AGI? Five contradictory answers that are all true at once — shaunralston · 2026-09-04
- Early Access Tester: GPT-6 Astra Built a Blender Werewolf in 8 Minutes and a World Simulator in 17 — TheMoonMidas · 2026-09-04
- swyx Burned 20B Tokens Stress-Testing Astra on Real AI Engineering Tasks — All for Under $6/Hour — charliermarsh · 2026-09-04