ObviousBench ranks LLMs on trivial-for-human questions by cost — GPT-Luna takes 2 crowns
adamallcock · reddit · 2026-09-23
Redditor adamallcock introduces ObviousBench, a benchmark where every question is trivial for humans but hard for LLMs, so a perfect 100% is achievable. The real game is minimizing cost/tokens to hit a score tier (90%+, 95%+, 99%+).
Current results: GPT-Luna takes first in the Max and Medium categories and third in Low. Leaderboard at obviousbench.com; paper published on Zenodo with a DOI.
More from Models
- Jev-style models: classifiers or "decision models"? The naming debate — MaziyarPanahi · 2026-09-23
- OpenAI launches GPT-6 Sol and Luna: Sol beats Claude Opus 5 at 9% of the cost — whoiskatrin · 2026-09-23
- "Opus 5.5 makes the Claude subscription worth it again" — TheOyinbooke · 2026-09-23
- Qwen 3.8 27B local coding session runs 3 days on one RTX 4090, then spews endless slashes — Tiny-Entertainer-346 · 2026-09-23
- Claude adds banked usage limit reset button on web and desktop — airesearch12 · 2026-09-23
- Bug Hunt Bench ranks GPT-6 Astra top as coding models fix real planted bugs, costs spread 200x — PawelHuryn · 2026-09-23