OpenAI's ~400 math results were a side effect of benchmarking its internal model

RexDouglass · x · 2026-10-10

deredleritt3r explains why OpenAI released only 400 math results: the exercise wasn't meant to "destroy math research" but to benchmark its internal model, which turned out to need nothing less than the hardest open problems. That's why all non-Millennium-Prize solutions came from a single agent after 3 hours of thinking. The solutions were a side effect of benchmarking, and OpenAI chose to publish them as genuine advances.

Andrew Curran's quoted context adds: most results weren't solved by a swarm like Navier–Stokes but one-shot by an internal model (nicknamed Aeon) from a single prompt. Aeon didn't exist before end of August, is still training, and its log-scale returns curve hadn't dropped off a month ago. Only 400 of 4000 problems were published — the rest likely solved since.

Original post →

More from Models

Models channel →