Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models
Kevin David Hayes, Arka Pal, Haosong Zhang, Tom Goldstein, Micah Goldblum
cs.AI
2026-09-22
An external calibrator reads only the question and response to score black-box API answers in one forward pass: 0.863 AUROC held-out, 0.814 zero-shot on 13 unseen models.
A physician asks GPT whether two drugs interact. The model answers confidently and wrongly, and nothing in the API response flags the failure. That is the everyday situation for teams building on closed-source models: GPT, Claude, and Gemini return no token log-probabilities and cannot be fine-tuned, yet these are exactly the models that most need a reliability signal in high-stakes deployments.
The academic toolbox does not fit. Logit- and hidden-state methods work but need white-box access. Black-box alternatives each pay a cost: verbalized confidence is systematically overconfident, self-evaluation inherits the model's blind spots, and semantic entropy multiplies inference cost by 5-10x through repeated sampling. Fast and accurate uncertainty for API models has been an open slot.
Pinocchio stops trying to extract uncertainty from the target model and trains an external judge instead. Given a question, the response, an optional model-identity tag, and any images, it outputs P(correct).
Pooling multiple target models is where transfer comes from. Single-source calibrators reach 0.672-0.759 mean AUROC on held-out targets; pooling three sources reaches 0.878. On an eleven-target transfer suite, four strong sources alone give 0.649, and adding the three weaker models in steps climbs to 0.814. Weaker models contribute the mistakes strong models no longer make.
On 1,953 held-out responses (question-level split, zero overlap):
| Method | AUROC | Brier | ECE |
| Pinocchio | 0.863 [.847, .879] | 0.164 | 0.107 |
| Combined (verbalized + length) | 0.649 | 0.229 | 0.021 |
| Verbalized (raw) | 0.610 | 0.326 | 0.299 |
| Verbalized (isotonic) | 0.644 | 0.231 | 0.009 |
| LLM-as-judge (GPT-5-mini) | 0.532 | 0.257 | 0.110 |
| TypeSafe Jev (commercial) | 0.666 | 0.240 | 0.133 |
All gaps pass DeLong's test (p < 0.001). Isotonic recalibration drives verbalized confidence's ECE to 0.009 but leaves AUROC at 0.644: recalibration fixes calibration, not ranking. Pinocchio's edge is discrimination.
Sampling methods fare worse when run faithfully. On LLaMA-3.1-8B with N=10 real samples from the target itself, semantic entropy (0.476), self-consistency (0.477), SPUQ (0.482), and a logit-free conformal method (0.574) all land at or below chance on a hard benchmark suite; Pinocchio, scored black-box on the same responses, gets 0.811. The pattern replicates on Granite-4.1-30B (0.806) and Devstral-2-24B (0.836). The explanation: on hard problems the model commits to one wrong trajectory across all samples, so consistency stops tracking correctness.
Zero-shot transfer: the same calibrator scores thirteen models it never saw, across eight organizations and including GPT-6 Astra, released after training, at 0.814 mean AUROC, only 0.05 below the in-distribution 0.862. Unseen benchmarks hold up too: 0.875 versus 0.877 in leave-K-out validation. On four open targets, isotonic recalibration on 100 labeled examples cuts ECE from 0.259 to 0.058.
The judge actually reads the answer. On 397 questions where the training-source models disagree (same question, same difficulty), it ranks correct responses above incorrect ones within a question at 0.726 AUROC, where any answer-blind predictor is pinned at 0.5. Stripping or injecting hedge words such as 'maybe' moves AUROC by less than 0.02, so it is not counting hedges.
Deployment workflows: error detection AUPRC of 0.867 versus 0.576 for verbalized confidence; confidence gating auto-executes 27-35% of responses at 90% accuracy, which no baseline reaches; ranking low-confidence responses for human review cuts the load 21% (79% instead of 100% of responses) to reach 95% system accuracy. Data scaling is log-linear: 2,000 examples give 91% of full performance and 5,000 give 96%, while calibrator size barely matters.
For teams shipping on closed APIs, this is the cheapest correctness signal available: one 0.8B forward pass against 5-10x inference for sampling methods, no access to logits, weights, or internals, and a two-line pip integration. The scores plug directly into gating, triage, and routing. The paper also demonstrates side uses: DPO pair selection (88.8% informative pairs), routing among three models (67.4% accuracy, +5.4 points over always using the single best model), and synthetic-data filtering (90.9% accuracy at 50% retention).
The method itself is incremental: it follows APRICOT's recipe of an auxiliary model trained on input-output pairs, with multi-target training, model-identity tags, vision-language signal, and demonstrated transfer to unseen models as the new parts.
One independent finding deserves attention: sampling-consistency methods collapse below chance on hard, long-form benchmarks, replicated across three target models. Teams using sampling for hallucination detection should re-examine that assumption.
Author-stated: the calibrator needs labeled data with automatically gradable answers, so open-ended generation is out of scope; cross-family transfer degrades (0.862 to 0.814); performance leans on benchmark and model-identity metadata that may be unknown at deployment; training data is English-only; it is weakest exactly on the hardest benchmarks (HLE-Multimodal 0.557), where correct responses hedge as much as wrong ones; and step-level uncertainty prediction is near random (0.53 AUROC).
Two open questions. Responses are truncated to 800 characters in training, so long chain-of-thought outputs are only partially visible to the judge; whether that loses signal on increasingly long reasoning outputs is not settled in the main text. And the healthcare and finance deployment scenarios price physician review at $25 per query and similar assumed values; those cost numbers are illustrative, not measured.