Debate: Model Frontiers and Erdős Problems
teortaxesTex · x · 2026-07-15
The quoted content debates how to view the model frontier: some argue that looking only at public coding benchmarks is insufficient because many good benchmarks are private, and samples are one-sided for judging capabilities. Even if some new models perform better on common leaderboards, it doesn't mean they've reached the true frontier.\n\nThe original poster retorts by pointing out that DeepSeek-V4, GLM-5.2, and Kimi-K2.6/2.7 haven't solved any Erdős problems yet, arguing that "not all Erdős problems are equally difficult," and encouraging further attempts. Overall, it uses math problems and benchmark limitations to emphasize that model capability evaluation shouldn't rely solely on standard leaderboards.
More from Models
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22
- Nanbeige4.2-3B launches as a 3B Looped Transformer model that beats larger baselines — Wooden-Deer-1276 · 2026-07-22
- A post says six companies now beat Google’s best LLM, including two open-source models — soham_btw · 2026-07-22
- Google says Gemini 3.5 Pro is in testing and Gemini 4 is already pre-training — Wide-Ad1564 · 2026-07-22
- Gemini 3.5 Flash Lite Tested: Not Frontier-Optimal, but Hits 350 tok/s — brandon_galang · 2026-07-22