Scoring 80+ on MMLU is Just Another Tuesday Now
andrew_n_carr · x · 2026-07-22
Tech professional Andrew Carr marvels at the rapid pace of the AI industry: just a year and a half ago, a large language model scoring over 80 on the MMLU (Massive Multitask Language Understanding) benchmark was an industry-shaking event; today, it's just another Tuesday.
This observation resonates with many, directly reflecting how quickly the baseline capabilities of frontier AI models are escalating, and how the psychological threshold for high-performing LLMs continues to rise.
More from AGI Musings
- AI lab staff have gone strangely quiet about next-year capability predictions — ChrisGPT · 2026-07-27
- Use an LLM to rank genuinely unsolved problems, not the ones already solved in training data — p_e_cooper · 2026-07-27
- AI slop is already clogging PR review and weakening the credit system behind science — rbhar90 · 2026-07-27
- AI could erode science by flooding research with credible slop — rbhar90 · 2026-07-27
- Organizations may already be the planet’s superintelligences — eldonredwards · 2026-07-27
- In the AI race, the only durable moats may be energy and information — GregKamradt · 2026-07-27