XConf Estimates LLM Confidence From the Model's Own Track Record, Beating Self-Consistency at 1/10 Cost
CambUni · hf · 2026-09-17
Existing confidence estimators only read the current inference process. XConf argues confidence should be grounded in the model's accumulated experience, stored as graded past episodes with reflections, stated confidence, outcomes, and lessons.
- Recall: retrieves similar past episodes with similar stated confidence and reads off their historical success rate
- Reflect: shows the model its record, has it name recurring failure modes, and restate confidence
The estimator needs no logit access or weight updates, costing one extra generation. Across nine benchmarks (reasoning, coding, multimodal QA, interactive agents) and four models, XConf matches or beats ten-sample self-consistency on AUROC in 23 of 24 comparisons, with far lower ECE, at a tenth of the cost. Abstaining on the 10% least-confident episodes lifts agent task success by up to 8.7 points.
More from Research
- Gowers responds to letter on maths and AI signed by 25 Fields medallists — tak3sh8 · 2026-09-17
- World models share one architecture — the tokenizer is where methods diverge — abursuc · 2026-09-17
- TMLR tightens desk rejects amid submission deluge, quizzes authors on their own papers — RexDouglass · 2026-09-17
- Anatomy of modern world models: a tokenizer compresses states, a module predicts the next — abursuc · 2026-09-17
- Speculative decoding: small draft model proposes tokens, big model verifies in one pass — HowDevelop · 2026-09-17
- Researcher proposes a journal for vibe-coded papers, with AI models as reviewers — peter_richtarik · 2026-09-17