XConf Estimates LLM Confidence From the Model's Own Track Record, Beating Self-Consistency at 1/10 Cost

CambUni · hf · 2026-09-17

Existing confidence estimators only read the current inference process. XConf argues confidence should be grounded in the model's accumulated experience, stored as graded past episodes with reflections, stated confidence, outcomes, and lessons.

The estimator needs no logit access or weight updates, costing one extra generation. Across nine benchmarks (reasoning, coding, multimodal QA, interactive agents) and four models, XConf matches or beats ten-sample self-consistency on AUROC in 23 of 24 comparisons, with far lower ECE, at a tenth of the cost. Abstaining on the 10% least-confident episodes lifts agent task success by up to 8.7 points.

Original post →

More from Research

Research channel →