ImageJevBench v0.2 launches on 573 images, Imajev 4B tops five-model ranking
airesearch12 · x · 2026-09-30
BenchmarkHeaven released ImageJevBench v0.2, an image-decision benchmark built on 573 fresh synthetic images with a public/sealed split, chance-corrected scoring, and separate computer and browser decision tracks. Imajev 4B ranks #1 in the five-model core ranking. The benchmark also publishes capability-vs-latency and capability-vs-cost trade-offs (USD per 1,000 decisions), a composite of intelligence, calibration, speed and cost, plus a new browser-use pilot measuring on-screen action. Selection chronology and an unscored 404 checkpoint are disclosed.
More from Models
- DepthBench aims to settle the race to beat the 'depth curse' in deep LLMs — FinanceYF5 · 2026-09-30
- User segments AI assistants by context: Claude for work, ChatGPT for everything else — enggirlfriend · 2026-09-30
- Benchmarker accuses AI model of 'dirty looping' after suspicious eval results — scaling01 · 2026-09-30
- Developer claims OpenAI silently degraded GPT-5.6 Sol performance weeks after release — Diveye · 2026-09-30
- DeepSeek open-sources full Ascend software stack, key benchmarks near hardware limits — lxfater · 2026-09-30
- 'We're no longer hiring motion designers': Opus 5.5 made a motion design video with just 3 prompts — FinanceYF5 · 2026-09-30