Jev trails Gemini and DeepSeek on calibration, but still handles more decisions solo
frappuccinoCoin · reddit · 2026-09-22
Jev Benchmarks measured Jev's confidence calibration; its training method is literally called "Reinforcement Learning for Calibrated Decisions." Calibration gap vs human labels (lower is better): yes/no Jev 5.0 vs Gemini 3.8 Flash 2.0; pick-one Jev 9.8 vs DeepSeek V4.1 Flash 2.8; rubric Jev 19.7 vs GLM-5.3 12.9.
Despite worse calibration, while holding 95% accuracy Jev still handles more decisions alone than any of them (86% of yes/no calls) — worse calibrated, but better at knowing when it's right.
More from Models
- Vals AI: Grok 4.7 drops to #24 on Vals Index, down 5 points from Grok 4.6 — zacharynado · 2026-09-22
- Grok 4.7 example: three-year financial analysis exposes currency-masked growth stall — ArtificialAnlys · 2026-09-22
- AA example: Grok 4.7 independently runs valuation chain and flags divergence from deal partner — ArtificialAnlys · 2026-09-22
- Grok 4.7 ranks just behind Anthropic's Opus 5 on AA-Briefcase at ~50% of the cost per task — ArtificialAnlys · 2026-09-22
- Replicating ExploitBench Would Cost ~$59.3M in API Fees, Security Researcher Estimates — OwariDa · 2026-09-22
- Xiaomi Releases Small Qwen 3.5 9B Distill SFT'd on MiMo Data, Plus RL Environments — teortaxesTex · 2026-09-22