ObviousBench ranks LLMs on trivial-for-human questions by cost — GPT-Luna takes 2 crowns

adamallcock · reddit · 2026-09-23

Redditor adamallcock introduces ObviousBench, a benchmark where every question is trivial for humans but hard for LLMs, so a perfect 100% is achievable. The real game is minimizing cost/tokens to hit a score tier (90%+, 95%+, 99%+).

Current results: GPT-Luna takes first in the Max and Medium categories and third in Low. Leaderboard at obviousbench.com; paper published on Zenodo with a DOI.

Original post →

More from Models

Models channel →