U-Space: Uncovering When and Why Uncertainty Arises in Language Models
Tobias Braun, Nils Loose, Alexander Herzog, Virginia Ceccatelli, Marcus Rohrbach, Thomas Eisenbarth, Lorenzo Cavallaro
cs.CL, cs.AI, cs.LG
2026-10-07
U-Lens scores a trace by four doubt directions times mean token entropy. Length-controlled AUROC beats the best baseline by 1.9-4.7 points on three models and four benchmarks.
Wrong answers from reasoning models tend to run long. Take the token count itself as the uncertainty score and, averaged over Gemma 4, Qwen3.5, Magistral 1.1, and four benchmarks, AUROC is already 68.6%. That is 0.2 points under Feature-Gaps, which is fit with correctness labels, and 1.0 point under TokUR, which spends five forward passes. Bin outputs into ten quantiles by length and score only inside each bin, and length falls to 52.4% AUROC. A confidence table has to show how much of it is just "longer means wronger."
The usual estimators are expensive, and they throw the location away. Semantic entropy draws ten samples and clusters them by meaning. TokUR uses five passes. P(True) asks the model to grade its own answer and, with no length control, reaches 56.7% AUROC. All of them end as one scalar. That scalar does not say which token brought the doubt in, or whether the doubt is vagueness, missing information, or evidence that disagrees with itself.
U-Space is a low-dimensional slice of the residual stream, the hidden-state channel every layer reads and writes. The slice is meant to hold doubt the model can put into words. Building it does not use correctness labels and does not update the model being read.
Anchors come from two published lists: 61 uncertainty cues from Chen et al., and 348 single-word certainty terms from Rocklage et al. rated at or above the corpus mean. Words the tokenizer splits are dropped. A separate sentence encoder then sorts the rest into ambiguity, incompleteness, conflicting evidence, and general uncertainty.
Each anchor is pulled from vocabulary space back into the residual stream with the J-Lens, a mean Jacobian. The Jacobian approximates how an intermediate state would move the final layer, so the model's own unembedding can read it. Gemma 4 and Qwen3.5 use public lenses. Magistral 1.1 has none, so the lens is estimated on 100 WikiText prompts at sequence length 128, by backpropagation through a frozen model. Within a category, the centroid of the uncertainty anchors minus the centroid of the certainty anchors cancels the shared confident register. Orthogonalizing the four axes yields the basis at that layer.
The U-Lens projects each reasoning token onto those axes, normalizes the coordinates so only the direction remains, and maps them through a softplus into the positive cone. A negative alignment on one category cannot cancel a positive alignment on another, so several kinds of doubt can be present together. The length of that vector is Acone. It is read once, at the end-of-thinking marker. That state has already conditioned on the full trace, at a fixed position, so the score does not pool a variable number of tokens.
The first-order factor is mean predictive entropy over the reasoning tokens: how flat the next-token distribution is. The second-order factor is Acone: whether the doubt lines up with one of the four verbal categories. Those names refer to levels of confidence, not to derivatives. The trace score is their product, with no learned mixing weight. Multiplying either factor by a positive constant leaves the ranking unchanged. The readout layer is locked at two-thirds depth: layer 40 of 60 in Gemma 4, 42 of 64 in Qwen3.5, and 27 of 40 in Magistral.
MMLU-Pro, Omni-MATH, SuperGPQA, and TriviaQA each contribute a fixed sample of 1,000 items. Three seeds, one response per item. Accuracy is 78.8% for Gemma 4, 79.7% for Qwen3.5, and 65.0% for Magistral 1.1. AUROC and AUPRC ask whether wrong answers sort ahead of right ones. AURC is the error that remains as the least certain answers are dropped. Lower is better.
With no length control, U-Lens averages 71.1±4.7 AUROC, 45.5±4.2 AUPRC, and 16.2±8.3 AURC. TokUR sits at 69.6 AUROC and the same 16.2 AURC. Length alone is 68.6. Maximum softmax probability is 53.3.
After length matching, U-Lens is the only method that ranks first on all nine model-by-metric comparisons. Against the strongest baseline on each model, AUROC is higher by 2.2, 1.9, and 4.7 points. The paper reports AUPRC gains of 0.8 to 4.2 points and AURC reductions of 0.4 to 2.5. U-Lens itself drops only 1.8 points from 71.1.
| Model | U-Lens | Best baseline | Predictive entropy | Acone only | Length |
| Gemma 4 | 68.4 | 66.2 Feature-Gaps | 66.1 | 66.3 | 52.1 |
| Qwen3.5 | 71.0 | 69.1 DeepConf / Self-Certainty | 68.9 | 66.7 | 52.7 |
| Magistral 1.1 | 68.3 | 63.6 DeepConf | 62.6 | 67.6 | 52.3 |
Which factor carries the score depends on the model. On Qwen3.5, entropy is stronger (68.9 versus 66.7 for Acone). On Magistral, Acone is stronger (67.6 versus 62.6). On Gemma 4 the two trade places across metrics. The product beats either factor alone. Against mean predictive entropy, the AUROC gaps are 2.3, 2.1, and 5.7.
The readout is not a keyword counter. Uncertainty cues appear in 23% to 34% of Gemma 4 traces, and in 4% of TriviaQA traces. On the 77% of Gemma 4 traces with no cue in the thinking block, length-matched AUROC is 68.6%, against 68.4% on the full set. Counting cue words falls well short of Acone once length is matched. Subtracting certainty-word counts pushes that lexical score below chance.
When Qwen3.5 answers a question about Hamlet, the trace is read first as incompleteness, then ambiguity, then conflicting evidence, and finally general uncertainty. Adding the ambiguity direction during reasoning, scaled as a fraction of the hidden state's own norm, makes the model check for a trick and then still answer 4. Steering each category during prefill turns "What is 2+2?" into talk of bases or jokes. The paper shows these traces. It does not report an accuracy drop.
Swapping the anchors, the same construction reads emotion and reward hacking in the appendix. On the 5,427-comment GoEmotions test set, anger, sadness, fear, and negative affect land at 74.6, 76.8, 79.8, and 73.0 AUROC. The cheating axis scores 90.2 to 92.6 on held-out responses and 91.4 to 96.9 on held-out cheat methods, against 59.3 to 60.7 for length. It separates absurd or pretentious text and misses plausible fake citations.
The lens is computed once per checkpoint. Every later answer costs a single generation. The scalar can drive abstention, and the four coordinates say which kind of doubt a span carries. Before length control, U-Lens is only 1.5 AUROC above TokUR, and AURC ties. After length control, semantic entropy leads mean predictive entropy by 0.1 points on the cross-model average. The margin over entropy is what is left once trace length is not allowed to do the sorting.
An API that hides the residual stream cannot build this basis.
This is not a calibrated safety detector. High-stakes use still needs an independent check. Steering the same directions changes how firmly the model commits.
The category directions are separated and far from orthogonal. Pairwise cosines among the ambiguity, conflicting-evidence, and incompleteness difference vectors are 0.51, 0.60, and 0.65. The certainty list mixes in words that only mark a confident manner of speaking. Anchors are English, and a multilingual set is left for later. An external encoder assigns words to categories. "No training" means the tested model's weights stay frozen.
Acone moves sharply with depth, so the headline numbers are tied to the two-thirds layer and to an end-of-thinking token. Chat models that do not emit a reasoning trace are not tested. Truncated and unparseable outputs are dropped. The main tables average four benchmarks, so they do not show which tasks fail.
On Magistral, unadjusted U-Lens AUROC is 65.7%, under length at 67.8% and under Acone at 68.5%. After length matching, the product is only 0.7 points above Acone (68.3 versus 67.6). Before length control, the cross-model standard deviation is ±4.7 AUROC and ±8.3 AURC. Steering shows that a direction can change the wording. It does not give a calibrated effect size. The emotion poles are twelve hand-picked words, tested on one model.