Frontier AI models begin crossing the human baseline on SimpleBench
Bojackin_Around · reddit · 2026-09-06
- SimpleBench targets reasoning humans find easy but LLMs have long struggled with: commonsense, spatial/temporal reasoning, social judgment, and traps for shallow pattern-matching.
- Plotted model scores are averages over 5 runs; random guessing scores 16.7% — close to where models started just over two years ago.
- The human baseline comes from 9 native-English, non-specialist participants, each answering a random subset of 25 questions, with at least high-school math proficiency (likely above average).
- The poster notes frontier models' average scores are now crossing this human baseline, a milestone on questions designed to resist shortcut learning.
More from Models
- Users report GPT-6 Astra finding up to 176x code speedups — or nothing at all — ivan_bezdomny · 2026-09-06
- Blogger: Chinese labs have cracked scaling and RL, need ~6 months to reach Astra level — zephyr_z9 · 2026-09-06
- Months of math work done in 26 minutes: Astra delivers 38-page constant-size proof — kfountou · 2026-09-06
- Hands-on: Astra one-shots a single-file Minecraft game, full sim done in 145 minutes — tegridyblues · 2026-09-06
- Compute likely tied up in pretraining, but Claude dev experience keeps getting worse — ATTlKA · 2026-09-06
- Normies can't tell Gemma 4 26B A4B from SOTA models in casual testing — TheMoonMidas · 2026-09-06