ConfTuner: Tokenized Brier Score Cuts Verbal Confidence ECE by up to 54.7%

ConfTuner: Training Large Language Models to Express Their Confidence Verbally

Yibo Li, Miao Xiong, Jiaying Wu, Bryan Hooi

NeurIPS 2025

cs.CL, cs.AI

2025-08-26

ConfTuner trains verbal confidence via tokenized Brier score, no proxy labels. ECE drops up to 54.7% vs the best baseline; AUROC +14.4%, helping self-correction and cascade.

What problem this solves

LLMs already emit percent-style confidence in text, but training never treats that number as a prediction that can be wrong. They still stamp 99% on errors. In medicine or law, people act on the stamp.

Prompting barely moves calibration. Fine-tuning has no ground-truth confidence, so earlier work uses proxies: group-level accuracy, agreement across samples, or a judge model. Group stats collapse hard and easy items into one score. Sampling is slow and noisy. A judge writes its bias into the labels.

ConfTuner asks whether verbal confidence can be trained with neither true confidence labels nor those proxies.

Method

ConfTuner, from the National University of Singapore, takes a proper scoring rule used for classifiers and applies it to tokens. Brier score (y − p)² is smallest when p equals the true chance of being correct. An LLM does not emit a scalar p. It emits a token sequence such as Confidence: 80%.

Figure 2 uses a grade-school check: is 1051 larger than 1039+15? The base model answers Yes at 99% confidence. After tuning, that same Yes comes with 5%. 1051 is less than 1054.

Two steps.

Nobody labels a target like 67% for this item. The same question is not sampled 10 or 100 times. The proof is that this loss is a proper scoring rule for verbalized confidence: the minimizer puts all mass on the token closest to the true correctness probability η.

Training data is HotpotQA only. Bases: Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Ministral-8B-Instruct-2410. Inference is greedy decoding.

Results

Comparisons are the untouched base, Ensemble (three samples averaged), and two trained methods, SaySelf and LACIE, re-run on the same bases and HotpotQA. Metrics: ECE (lower better) and AUROC (higher better).

BaseMethodECE mean (5 sets)AUROC mean (5 sets)
LLaMABase0.27680.5923
LLaMALACIE0.23870.6229
LLaMAConfTuner0.10820.6740
QwenBase0.37810.6155
QwenConfTuner0.28720.6861
MinistralBase0.43930.5216
MinistralConfTuner0.18840.6810

Versus the strongest baseline, ECE drops by up to 54.7% (LLaMA against LACIE) and AUROC rises by up to 14.4% (Ministral against Ensemble). On in-distribution HotpotQA, LLaMA ECE falls from 0.4803 to 0.0405. TruthfulQA and TriviaQA improve as well. GSM8K and StrategyQA move less; Qwen GSM8K ECE goes from 0.1306 to 0.1302.

Averages win. Some cells do not. Qwen StrategyQA ECE is 0.1815, behind Ensemble at 0.1226. LLaMA TruthfulQA AUROC is 0.5739, behind Ensemble at 0.6038. Ministral StrategyQA AUROC is 0.5147 against Ensemble 0.6222.

Trained on numeric scores and tested on high/medium/low wording, LLaMA AUROC still averages 0.6511 versus SaySelf 0.5989. Asked to express uncertainty in free text, with GPT-4o mapping the phrasing onto 0 to 100, implicit ECE averages 0.1483 and AUROC 0.6820, against explicit 0.1082 and 0.6740.

A LLaMA ConfTuner that scores GPT-4o answers cuts GPT-4o's self-reported mean ECE from 0.1640 to 0.0970 and lifts AUROC from 0.5945 to 0.6275.

Self-correction rewrites answers below 0.5 confidence. Qwen plus ConfTuner shows the largest accuracy gains on HotpotQA and TruthfulQA. Baselines move little or get worse, because they mark correct answers as low-confidence. In a cascade, 100 to 400 lowest-confidence items go to GPT-4o. Under the same budget, accuracy is up to 9.3 points higher on HotpotQA and 5.5 on TruthfulQA.

Training takes 4 minutes on 4×A40, 2,000 examples, one sample per question. SaySelf: 120 minutes, 90,000 examples, 100 samples each. LACIE: 26 minutes, 10,000 examples, 10 samples. The appendix says 2,000 examples already saturate.

Why it matters

The recipe is small: gold answers are enough. No confidence labels, no sampling to build a target. Users read verbal confidence. Closed APIs do not expose logits.

Two uses are close to production. Add a confidence line to a 7B or 8B instruct model. Or hang an external calibrator on GPT-4o. Self-correction and cascade only pay off after calibration; otherwise rewrite budget lands on items that were already right.

On Qwen, mean ECE only moves from 0.3781 to 0.2872. The three bases do not move equally. The loss never targets answer accuracy.

Limitations

The authors flag two gaps. A fixed confidence token set may not stretch to richer dialogue or more varied hedges. A proper scoring rule sets the right objective; data quality, architecture, and optimizer dynamics still decide whether calibration shows up.

The main text never prints the accuracy table in Appendix F, so any hit to answer quality cannot be checked from the body. Self-correction is a change-rate bar chart, with no before/after absolute accuracy. Implicit confidence is scored by GPT-4o, a model that is itself overconfident. Training is HotpotQA only; Qwen's GSM8K ECE barely moves, so transfer is uneven. Tokenized Brier conditions on the greedy answer's correctness. Whether η holds under temperature sampling or another decoder is untested.

Terms

Source

What people are saying

Related papers

All paper explainers