ImageJevBench: image decision benchmark ranks top models, full eval costs $0.02
airesearch12 · x · 2026-09-26
A new image decision benchmark, ImageJevBench v0.1, is live on Benchmark Heaven, evaluated over 89 public synthetic images:
- Top 3: 🥇 Jev-Omni, 🥈 decider-2b-vision, 🥉 Reflex 4B
- Running the top 5 models across all images cost just $0.0217, which the author calls disruptive for eval economics
- Benchmark Heaven also aggregates model comparisons with filters for hosting region, data confidentiality, and pricing bases
More from Models
- User generates an anime-style fight scene entirely with code using Claude Opus 5.5 — EricBuess · 2026-09-26
- Dev verdict on Opus 5.5: the first good Opus since 4.8 — holdenmatt · 2026-09-26
- Opus 5.5 code review demo draws attention — tlakomy · 2026-09-26
- One prompt, zero libraries: Opus 5.5 builds polished vanilla JS motion graphics in Claude Code — EricBuess · 2026-09-26
- "Jev was in fact not beaten": open-source Jevons paradox claim walked back — BLUECOW009 · 2026-09-26
- Developer Gives a Persistent AI Instance Its Own Twitter Account to Post Autonomously — repligate · 2026-09-26