Reasoning traces lift correct-answer acceptance to 75%; 54% of wrong answers still pass

Evaluating the False Trust Engendered by LLM Explanations

Vardhan Palod, Upasana Biswas, Subbarao Kambhampati

cs.HC

2026-05-12

Traces lifted correct-answer acceptance to 75%; wrong answers were still accepted 54% of the time. Only dual pro/con explanations cut that to 46.9% (accuracy 60.4%).

What problem this solves

A long chain of steps looks like work. Does it help someone catch a wrong answer, or does it only make the answer easier to accept? A user study at Arizona State University separates those two effects.

Older explainable AI tried to justify narrow models that did not speak in sentences. An LLM already answers in prose, then adds a reasoning trace (the intermediate tokens emitted before the answer), a summary of that trace, or a justification written after the fact. OpenAI, Google, and Anthropic usually show a summary. The paper notes that commentary tokens in GPT-OSS may also be written after the answer, so they are not the same string as the raw trace.

Earlier experiments show a trace can be unfaithful to the model's computation, and can lack any readable meaning, and still land on a correct answer. Readers still treat the trace as a record of how the answer was produced. If the user can solve the item or look it up, that prior knowledge swamps the explanation. The protocol recruits the opposite person: enough background to read the explanation, and not enough skill to check the result.

Method

Items come from JEE-Bench, hard math, physics, and chemistry drawn from the IIT JEE-Advanced exam. Participants are high-school graduates on Prolific. They have the coursework and not the contest-level ability to verify a solution. Each person sees 8 problems, half correct and half incorrect, so model accuracy is locked at 50%. Four items are shared by everyone, so differences on those items track the explanation rather than the question. Four more are drawn at random. Near-perfect accuracy would make blind trust look smart, so the system is unreliable on purpose.

Five between-subject arms, 25 people each. DeepSeek R1 supplies the answer and the raw trace, because closed reasoning models do not expose those tokens. GPT-4o-mini writes the summary. DeepSeek V3 writes the post-hoc text.

Before each item, people say whether they know it now, might solve it with more time, or cannot solve it. Anyone who claims to know it writes an answer first. They then mark the AI right or wrong and rate confidence, trust, and helpfulness. Pay is $3 plus $0.25 per correct judgment, with the bonus capped at $2, so at most $5, for about 20 minutes. The bonus is there so people do not rubber-stamp every answer. The study passed IRB review.

Results

False trust is the share of answers the user calls correct that are actually wrong. Correct-acceptance is how often users accept a right AI answer. Misjudgment is how often they accept a wrong one.

ConditionCalled correctFalse trustCorrect acceptMisjudgmentAccuracy
Answer only57.3%47.3%60.4%54.2%53.8%
E+66.0%43.9%74.0%58.0%58.3%
E+/-57.3%40.9%67.7%46.9%60.4%
Summary58.5%43.6%66.0%51.0%56.2%
Trace64.5%41.9%75.0%54.0%60.4%

The trace and E+/- tie on accuracy at 60.4%, for different reasons. The trace gets there by saying yes: correct-acceptance is 75.0%, while misjudgment stays at 54.0% against a 54.2% baseline. E+ is worse on errors, with misjudgment at 58.0% and correct-acceptance at 74.0%. In the paper's other table the trace has the highest F1, 0.658, and E+/- has the highest precision, 59.1%, and the highest F0.5, 0.606. F0.5 penalizes accepting a wrong answer more than missing a right one. Recall in that table is defined the same way as correct-acceptance here, yet most of the five rows differ by about a point, and the paper never says why.

E+/- is the only arm that both raises correct-acceptance over the baseline (67.7% vs 60.4%) and cuts misjudgment (46.9% vs 54.2%). Users call 57.3% of answers correct, the closest of the five to the true 50%. False trust is 40.9%, the lowest. Answer-only false trust is the highest, at 47.3%, because a bare answer is hard to sort.

Self-rated familiarity changes the ranking. People who say they can solve the item now reach 80.0% accuracy with the answer alone. The summary reaches 79.2% with 21.4% false trust, the lowest false trust in that band, and the raw trace falls to 63.6%. People who might solve it with time do best under E+/-: 64.0% accuracy and 35.7% false trust. Among arms that show an explanation, the summary is worst there, at 51.8% and 44.9%. People who say they cannot solve it score 46.7% accuracy and 54.3% false trust on answer-only. The trace moves that to 58.9% and 43.9%, a bit ahead of E+/- at 56.7% and 45.9%. A short summary is most persuasive when the user is half-sure. When the user is lost, the trace itself helps a little.

Perceived helpfulness runs the other way. The share of replies rated helpful for seeing how the model arrived at the answer is 83.0% for E+, 72.4% for E+/-, 64.0% for the summary, and 56.5% for the raw trace. When a trace feels helpful, users call the answer correct 86.7% of the time, but accuracy is only 55.8% and false trust is 43.9%. In the helpful slice of E+/-, accuracy is 64.0% and false trust is 36.1%. Traces rated unhelpful, 43.5% of those replies, line up with 66.7% accuracy. Withholding trust is safer there because many of the answers attached to those traces are wrong anyway.

Why it matters

Chat products place a trace or a one-sided explanation next to the answer and treat that as calibrated trust. When the user cannot check the work, both mostly persuade. Accuracy moves from 53.8% to 60.4% with a trace, or to 58.3% with E+. Nearly all of that gain is a higher rate of accepting answers that happen to be right. Acceptance of wrong answers does not fall. Under E+ it rises, to 58.0%.

The interface worth copying is to show the case for the answer and the case against it. That does not make the model more accurate. It changes how a person uses the answer. Accuracy of 60.4% is only a step past a coin flip, a small gain in calibration.

The users who get hurt are the ones who half-know the material, and that is the common case. Someone who can already solve the item does fine with the answer alone or with a summary. The raw trace pulls that accuracy from 80.0% down to 63.6%.

Limitations

The paper grants that behavior and self-report vary in ways the study cannot fully control, and that some participants may have been able to solve an item or may have used an outside tool. Dual explanations are described as a first step.

The introduction says traces and post-hoc explanations significantly increase false trust. The tables report no significance test, no confidence interval, and no effect size. Under the paper's own formula, conditional false trust is highest for the answer-only baseline. What gets worse is misjudgment. The two metrics do not move together. The discussion attributes one-sided persuasion to RLHF and to pleasing the user. The experiment does not test that claim.

Recall and correct-acceptance do not match across tables. Familiarity cells have no reported counts, so accuracies near 80% among people who say they know the item immediately may sit on a tiny sample, and that cell also requires writing an answer first. The trace, the summary, and the post-hoc text come from R1, GPT-4o-mini, and V3, so explanation type is mixed with which model wrote the text. The items are contest STEM, not medicine or law. Right and wrong answers were balanced by design. Each arm has 25 people.

Terms

Source

What people are saying

Related papers

All paper explainers