DeepSeek vs Jev: How an LLM Stacks Up on a System-One Probability Benchmark
frappuccinoCoin · reddit · 2026-09-22
A Reddit post compares DeepSeek against Jev on the jevals.com benchmark. Jev is a "System One" model — no text, just a probability per option — scored on both accuracy and honest calibration. LLMs can attempt it only by emitting probabilities with reasoning off. Result: DeepSeek trails Jev slightly on accuracy while being slower and pricier, but is the best-calibrated model on the board.
More from Models
- JevBench v1.3.0 launches: original Jev leads at 74.4 with 47 challengers closing in — airesearch12 · 2026-09-22
- Grok 4.7 Fast is the same model at 2x token rates, only in Cursor and Grok Build — Daniel_Farinax · 2026-09-22
- Commenter praises async 4 for disclosing training data mix percentages — stochasticchasm · 2026-09-22
- Grok 4.7 reportedly released as a fully agentic model built for Grok Bot — elonmusk · 2026-09-22
- OpenAI criticized for claiming 100 open math problems solved without disclosing the total attempted — burny_tech · 2026-09-22
- Matthew Berman Reviews Grok 4.7: 'I Don't Know How to Feel About It' — Matthew Berman · 2026-09-22