Sys1Cal-v1 Dataset Finds Jev Suppresses a Third Truth Value, Boosting Soft Accuracy to 0.978
Riccardo Porcedda · hf · 2026-09-30
The author introduces Sys1Cal-v1, a dataset for testing probability calibration of "System One Models"—foundation models that return structured decisions with probability distributions instead of text.
Motivation: Jev's central claim that its probabilities are calibrated lacks public testing; existing benchmarks measure confidence calibration, not whether each option's probability has the correct numerical meaning.
Method: Each item is a True/False question where P(A) is known by construction, queried via Jev's three primitives (Noul, Choice, Score) and scored by total variation distance from the ground-truth distribution. Open-source Choice-style baseline SemIf is also evaluated.
Key finding: In Choice answers, P(A) and P(neg A) sum to 1, but a term P(U)≠0 is missing—Jev suppresses a third truth value beyond True and False. Recovering P(U) improves median soft accuracy from 0.771 to 0.978, suggesting the model wants to say "I don't know" even on binary decisions.
More from Research
- Video lecture series by Stephen Wright, Yousef Saad and Peter Bartlett on ML optimization now available — caglar_ee · 2026-09-30
- Arbor: open-source framework for AI agents doing autonomous long-horizon research — burkov · 2026-09-30
- Physics-aware losses keep grain boundaries real when AI generates alloy microstructures — bravo_abad · 2026-09-30
- SOSP26 Paper YoloFS Targets Agent Filesystem Misuse, Built From 290 Real Incident Reports — tianyin_xu · 2026-09-30
- Agents can delete their own logs: Claude Code, Codex, others fail trace integrity, paper finds — maksym_andr · 2026-09-30
- NUS Proposes StoryEngine: A State-Grounded Agentic Framework for Coherent Long-Form Video Storytelling — NationalUniversityofSingapore · 2026-09-30