JevBench: DeepSeek V4.1 Flash outscores leader at 1/15th the cost per decision
airesearch12 · x · 2026-09-24
Benchmark Heaven's JevBench shows DeepSeek V4.1 Flash (thinking) scoring 97.8% on public items vs 86.6% for top-ranked Jev 1.13.0, while costing $0.040 vs $0.594 per 1,000 decisions — roughly 15x cheaper. The site offers filters for cost basis, EU hosting, and data confidentiality; the poster argues cheap Jev-class models are disruptive.
More from Models
- FLock's THIS/THAT 1.2 decision model beats Claude Opus 5 and GPT-5.6 with one forward pass — matlabulous · 2026-09-24
- Fable can now interrupt itself mid-task to answer a second prompt, then resume the first — gleech · 2026-09-24
- Q Labs: LLMs are depth-bottlenecked, loss keeps improving to 128 layers — rickasaurus · 2026-09-24
- Opus-5.5 is 2-3x faster and 60% cheaper than Astra, dev says in hands-on — haydendevs · 2026-09-24
- New AI models launch agents-first while chat becomes an afterthought: GPT-6 Sol missing from ChatGPT — mark_k · 2026-09-24
- Dev tells Sam Altman he's 'already Opus pilled' amid model loyalty banter — haydendevs · 2026-09-24