Paper: Calibration Should Be a First-Class Criterion in LLM Evaluation

Mario Sanz-Guerrero · hf · 2026-09-24

This paper argues calibration — the alignment between a model's expressed confidence and empirical correctness — should be treated as an essential property of every LLM, not a niche NLP topic. The problem is adoption, not measurement: new models, datasets, and benchmarks regularly ship without checking whether confidence scores are meaningful.

The authors call for each NLP subfield to pair its main metric with a calibration score.

Original post →

More from Models

Models channel →