Debate: Model Frontiers and Erdős Problems

teortaxesTex · x · 2026-07-15

The quoted content debates how to view the model frontier: some argue that looking only at public coding benchmarks is insufficient because many good benchmarks are private, and samples are one-sided for judging capabilities. Even if some new models perform better on common leaderboards, it doesn't mean they've reached the true frontier.\n\nThe original poster retorts by pointing out that DeepSeek-V4, GLM-5.2, and Kimi-K2.6/2.7 haven't solved any Erdős problems yet, arguing that "not all Erdős problems are equally difficult," and encouraging further attempts. Overall, it uses math problems and benchmark limitations to emphasize that model capability evaluation shouldn't rely solely on standard leaderboards.

Original post →

More from Models

Models channel →