OpenAI publishes 722 math papers averaging 3 hours of compute, claiming 90 of top 500 open problems
Latent Space · rss · 2026-10-07
OpenAI releases 722 math manuscripts from an unreleased internal model
- 722 manuscripts across 372 result families, drawn from 4,000 research problems; average 3 hours of ChatGPT Pro thinking compute per result. Altman calls it "a new era of discovery".
- Claimed (unverified) highlights: a quasi-Riemann Hypothesis result ("biggest number theory result in 200 years"), integer multiplication faster than n log n, uniqueness for the 3D elastic inverse problem open since 1994, partial progress on Riemann/Hodge/BSD. Mathematician Levent Alpöge hails it as "the most significant moment in mathematical history", while noting scooping and conflict-of-interest concerns.
- Skeptics: Will Depue built citedbyagi.com to track citations; an analysis finds 20% of results are disproofs/counterexamples, undercutting the brute-force narrative; Chollet asks whether gains generalize beyond RLVR-friendly domains.
Mistral Large 4 "Le Chonk": 1T total / 49B active params, natively multimodal, trained on 3,800 GB300s in Europe; $1.36/$4.18 per M tokens, open weights promised end of October. Beats GLM 5.3 in human evals, #2 in blind coding review behind Opus 5; AA Intelligence Index 38 (level with GPT-6 Luna) but 4x cost of similar-intelligence open models; Cline credits its cyber lead partly to fewer refusals.
Other releases: Google EmbeddingGemma 2, the first natively multimodal open embedding model (Apache 2.0, 270M–740M variants); Nano Banana 2.1 at $0.034/image; decision models as a category — OpenAI Decisions API beta (claims 10x faster) and Perplexity's open-weight pplx-decider; Ling 3.1 Flash (560B/25B active, AA index 41).
More from Models
- _xjdr names the only four open-weight models he finds interesting right now — _xjdr · 2026-10-07
- Researcher _xjdr Praises Inkling Architecture, Calls DSV4 Models an Acquired Taste — _xjdr · 2026-10-07
- Dev: OpenAI Decisions API needed fallback on 65%+ of requests, misjudged statue — taufiqintech · 2026-10-07
- Users still tag @grok on X even though the feature no longer works — BLUECOW009 · 2026-10-07
- "Opus 5.5 just knows": Users impressed by the model's intent understanding — rudrank · 2026-10-07
- Teknium pushes back on benchmark criticism: half of Hermes users use it for coding — Teknium · 2026-10-07