Jev trails Gemini and DeepSeek on calibration, but still handles more decisions solo

frappuccinoCoin · reddit · 2026-09-22

Jev Benchmarks measured Jev's confidence calibration; its training method is literally called "Reinforcement Learning for Calibrated Decisions." Calibration gap vs human labels (lower is better): yes/no Jev 5.0 vs Gemini 3.8 Flash 2.0; pick-one Jev 9.8 vs DeepSeek V4.1 Flash 2.8; rubric Jev 19.7 vs GLM-5.3 12.9.

Despite worse calibration, while holding 95% accuracy Jev still handles more decisions alone than any of them (86% of yes/no calls) — worse calibrated, but better at knowing when it's right.

Original post →

More from Models

Models channel →