Burned 300B Tokens with Nothing; Internal Model Broke Through in 3 Hours
burny_tech · x · 2026-10-08
Peter Gostev of Moonshot AI weighed in on the debate over the gap between OpenAI's internal model and Astra with a concrete case: on the Hadwiger–Nelson problem he burned roughly 300 billion tokens (GPT-5.6 and Astra) and got nowhere, while Bell reportedly achieved a breakthrough with the internal Pro model in about 3 hours (ruling out five colours and pushing the lower bound to 6 or 7 — not a full solution). He asked Astra to compare his existing work against the paper, and it concluded he was "two mathematical breakthroughs away." To him the capability gap feels impossibly large, though he notes it's a single example. This contrasts with ConstantinKogl, who argued the internal model is better but not dramatically so, having reproduced the paper-148 result himself with Astra in a 6-hour session.
Related event: OpenAI Internal Model Cracks Math Problem After 300B Tokens Spent(2 posts)→
More from Models
- Dev debate: is ColBERT-style late interaction still a cross-encoder as rerankers fade? — CShorten30 · 2026-10-08
- Claude Projects quietly adds scheduled tasks for automated recurring runs — ColleenMBrady · 2026-10-08
- Leak: X preps all-in-one subscription bundling X, Grok and Cursor in one usage pool — nima_owji · 2026-10-08
- 4 models, one two-file bug: 3/4 passed, 10x cost spread, and the cheapest run was the failure — lulzxdxdxd · 2026-10-08
- Apollo Research: Final-Checkpoint Evals Can't Catch Misalignment That Emerges Early — dl_weekly · 2026-10-08
- AI spend pulled both ways: cheap models cut bills while video generation drives them up — c_valenzuelab · 2026-10-08