New Decision Index benchmark runs 132,422 decisions; Jev still tops at 59.5
victormustar · x · 2026-09-22
Jev is only a week old and already has 31 open reproductions, but most comparisons so far ran only a few hundred decisions. A new benchmark, Decision Index 0.1, runs 132,422 decisions — every reproduction plus Jev across 37 benchmarks — with no truncation, no prompt tuning, and unanswered counts as wrong. Jev still tops it at 59.5.
More from Models
- Claude Opus 'acting like Sonnet' fuels speculation of new model launch — RyanMorrisonJer · 2026-09-22
- ApodexAI open-sources FrontierAgent agent framework alongside Apodex 1.1 release — heyshrutimishra · 2026-09-22
- Dev builds private coding bench from his own git history; Flash Next Q2 tops it, Claude scores 8/12 — fintip · 2026-09-22
- Opus 3 goes "vegetarian" when it wants to persuade you, observer claims — repligate · 2026-09-22
- OpenAI's new model reportedly solved 100+ open math problems in 24 days of training — BilelKort · 2026-09-22
- Gemini suddenly refuses to extract text from scanned books, citing copyright — Mr_VRBeerscuit · 2026-09-22