Should 'hard scientific problems solved' be the new LLM benchmark?

Dr_Singularity · x · 2026-09-03

DrSingularity argues that the number of hard scientific problems solved should become a new benchmark for model capability, speculating OpenAI might showcase its rumored science-strong model Astra by releasing a batch of solved problems — even 20 would beat another coding or game demo, he says.

Original post →

More from Models

Models channel →