Ericsson's MARGIN calibrates multi-agent LLM confidence at runtime, lifting selection accuracy by up to 14 points
EricssonAB · hf · 2026-10-09
Ericsson researchers propose MARGIN, a runtime confidence-calibration method for coordinating heterogeneous foundation models, whose self-reported confidences mean different things across responders and shifting workloads. MARGIN learns model-specific corrections from observed answer outcomes without retraining or a held-out calibration set.
Method: it tracks recent accuracy vs. stated confidence within confidence bands, corrects reported confidence by their ratio, and blends sparse-band corrections toward model-level estimates; corrected scores weight candidate answers in collective decisions.
Results: evaluated on code generation, QA, and math with an 18-model pool and a 9-model subset for distribution shift. On BigCodeBench, model-mean confidence is negatively correlated with accuracy, and picking the more confident responder in correct/incorrect pairs performs below chance. Against five online calibration baselines, MARGIN achieves lower post-shift ECE in most transitions, and boosts answer-selection accuracy by 4.3 and 14.0 percentage points on two of three code-generation benchmarks.
More from coding & agent
- Deepkit author: Bun is rediscovering our decade-old ideas, agents could revive it — MarcJSchmidt · 2026-10-09
- Qdrant ships a LangGraph memory store integration for semantic long-term agent memory — qdrant_engine · 2026-10-09
- Qdrant ships LangGraph memory store integration for vector-native agent memory — qdrant_engine · 2026-10-09
- Building a production MCP server: why DCR + PKCE-only is the auth combo that works — Actual_Tradition_990 · 2026-10-09
- gemini-cli PR adds canCreateSymlinks() check to skip failing Windows symlink tests — supunyasanthaofficial · 2026-10-09
- Hot-reload Lua on a virtual ESP32 and vibe-code hardware live from Claude — genmon · 2026-10-09