Jev model goes viral as NUS ConfTuner paper backs the same probability-calibration route

机器之心 · wechat · 2026-10-04

TypeSafeAI's Jev model is going viral: given a state and a question, it returns predefined options, scores, or event probabilities that code can consume directly, with emphasis on parallel probability prediction and calibration. Architecture details remain closed.

The approach echoes NeurIPS 2025 paper ConfTuner from NUS, whose Tokenized Brier Score trains the full distribution over candidate confidence tokens (0–100) read from logits — a mathematically proven Proper Scoring Rule that only needs correctness labels, no manual confidence annotation. Across five datasets, average ECE dropped from 0.2768/0.3781/0.4393 to 0.1082/0.2872/0.1884 for LLaMA/Qwen/Ministral; fine-tuning took 4 minutes on 4×A40 with 2,000 examples, versus 26–120 minutes for baselines. Calibrated probabilities enable model cascades, boosting HotpotQA and TruthfulQA accuracy by up to 9.3% and 5.5% at equal compute. The team also open-sourced JevTuner, an exploration of extending the idea from confidence tokens to business-decision tokens.

Original post →

More from Models

Models channel →