DeepSeek vs Jev: How an LLM Stacks Up on a System-One Probability Benchmark

frappuccinoCoin · reddit · 2026-09-22

A Reddit post compares DeepSeek against Jev on the jevals.com benchmark. Jev is a "System One" model — no text, just a probability per option — scored on both accuracy and honest calibration. LLMs can attempt it only by emitting probabilities with reasoning off. Result: DeepSeek trails Jev slightly on accuracy while being slower and pricier, but is the best-calibrated model on the board.

Related event: Jev's text-free model leads decision share but lags in calibration vs Gemini, DeepSeek(2 posts)→

Original post →

More from Models

Models channel →