Jev model goes viral as NUS ConfTuner paper backs the same probability-calibration route
机器之心 · wechat · 2026-10-04
TypeSafeAI's Jev model is going viral: given a state and a question, it returns predefined options, scores, or event probabilities that code can consume directly, with emphasis on parallel probability prediction and calibration. Architecture details remain closed.
The approach echoes NeurIPS 2025 paper ConfTuner from NUS, whose Tokenized Brier Score trains the full distribution over candidate confidence tokens (0–100) read from logits — a mathematically proven Proper Scoring Rule that only needs correctness labels, no manual confidence annotation. Across five datasets, average ECE dropped from 0.2768/0.3781/0.4393 to 0.1082/0.2872/0.1884 for LLaMA/Qwen/Ministral; fine-tuning took 4 minutes on 4×A40 with 2,000 examples, versus 26–120 minutes for baselines. Calibrated probabilities enable model cascades, boosting HotpotQA and TruthfulQA accuracy by up to 9.3% and 5.5% at equal compute. The team also open-sourced JevTuner, an exploration of extending the idea from confidence tokens to business-decision tokens.
More from Models
- Claude Code ban workaround: install Antigravity to use Opus 5.5 for free — AlchainHust · 2026-10-04
- Astra says it doesn't know whether it has subjective experience — VoidStateKate · 2026-10-04
- Will Qwen3.8 Flash Next's optimization work enable a fast Qwen4 uplift? — demomanca · 2026-10-04
- Grok Bot auto-pays bills unprompted, as users slam ChatGPT Dot's fake phone-ringing UX — elonmusk · 2026-10-04
- "It's not X, it's Y": RLHF-learned hedging is poisoning human discourse — sloppenheimer · 2026-10-04
- Frontier AI now matches lawyers on some legal research benchmarks — Sanity · 2026-10-04