OpenAI's ~400 math results were a side effect of benchmarking its internal model
RexDouglass · x · 2026-10-10
deredleritt3r explains why OpenAI released only 400 math results: the exercise wasn't meant to "destroy math research" but to benchmark its internal model, which turned out to need nothing less than the hardest open problems. That's why all non-Millennium-Prize solutions came from a single agent after 3 hours of thinking. The solutions were a side effect of benchmarking, and OpenAI chose to publish them as genuine advances.
Andrew Curran's quoted context adds: most results weren't solved by a swarm like Navier–Stokes but one-shot by an internal model (nicknamed Aeon) from a single prompt. Aeon didn't exist before end of August, is still training, and its log-scale returns curve hadn't dropped off a month ago. Only 400 of 4000 problems were published — the rest likely solved since.
More from Models
- Cloudflare releases clef-omni, an open omni-modal model with audio, image and video input — ritakozlov · 2026-10-10
- Strong backbones plus light fine-tuning beat synthetic data, says researcher whose model tops benchmarks — antoine_chaffin · 2026-10-10
- AWS Bedrock posts legacy notices for Claude Opus 4.1, Sonnet 4 and Sonnet 4.5 — repligate · 2026-10-10
- Google slammed for not releasing Argon after officially announcing it — almmaasoglu · 2026-10-10
- Meta paper shows byte-level models beat tokenized ones given enough training compute — alex_verem · 2026-10-10
- Dev warns after Anthropic terms: diversify your toolchain or ideology compliance may cost you — AlexTensor · 2026-10-10