Sys1Cal-v1 Dataset Finds Jev Suppresses a Third Truth Value, Boosting Soft Accuracy to 0.978

Riccardo Porcedda · hf · 2026-09-30

The author introduces Sys1Cal-v1, a dataset for testing probability calibration of "System One Models"—foundation models that return structured decisions with probability distributions instead of text.

Motivation: Jev's central claim that its probabilities are calibrated lacks public testing; existing benchmarks measure confidence calibration, not whether each option's probability has the correct numerical meaning.

Method: Each item is a True/False question where P(A) is known by construction, queried via Jev's three primitives (Noul, Choice, Score) and scored by total variation distance from the ground-truth distribution. Open-source Choice-style baseline SemIf is also evaluated.

Key finding: In Choice answers, P(A) and P(neg A) sum to 1, but a term P(U)≠0 is missing—Jev suppresses a third truth value beyond True and False. Recovering P(U) improves median soft accuracy from 0.771 to 0.978, suggesting the model wants to say "I don't know" even on binary decisions.

Original post →

More from Research

Research channel →