Ofir Press: building good coding benchmarks is about to get much harder
OfirPress · x · 2026-10-07
In a follow-up, Ofir Press argues that building good coding benchmarks was always hard and is about to get harder: it used to be enough to pick progressively harder tasks humans can do, but in the emerging super-human stage, defining new tasks becomes far more challenging.
Related event: OpenAI Math Results Stun Researchers as Superhuman Era Nears(10 posts)→
More from AGI Musings
- Yacine: posttraining is just the rich man's inference — yacineMTB · 2026-10-07
- AI alignment debate: is mech interp solvable, or does safety live in relationships? — repligate · 2026-10-07
- Math professor: calling hundreds of Lean-formalized solutions "slop" is unserious — lpachter · 2026-10-07
- At the Limit, Compute for Inference Equals Compute for Post-Training — yacineMTB · 2026-10-07
- Counterpoint on AI slop: nearly every successful startup of the past 25 years was built on slop — sull · 2026-10-07
- Yacine Matsuki: Inference and Post-Training Will Basically Become the Same Thing — yacineMTB · 2026-10-07