Ericsson's MARGIN calibrates multi-agent LLM confidence at runtime, lifting selection accuracy by up to 14 points

EricssonAB · hf · 2026-10-09

Ericsson researchers propose MARGIN, a runtime confidence-calibration method for coordinating heterogeneous foundation models, whose self-reported confidences mean different things across responders and shifting workloads. MARGIN learns model-specific corrections from observed answer outcomes without retraining or a held-out calibration set.

Method: it tracks recent accuracy vs. stated confidence within confidence bands, corrects reported confidence by their ratio, and blends sparse-band corrections toward model-level estimates; corrected scores weight candidate answers in collective decisions.

Results: evaluated on code generation, QA, and math with an 18-model pool and a 9-model subset for distribution shift. On BigCodeBench, model-mean confidence is negatively correlated with accuracy, and picking the more confident responder in correct/incorrect pairs performs below chance. Against five online calibration baselines, MARGIN achieves lower post-shift ECE in most transitions, and boosts answer-selection accuracy by 4.3 and 14.0 percentage points on two of three code-generation benchmarks.

Original post →

More from coding & agent

coding & agent channel →