Four LLM tutors intervene on 90% of problems at relative time 0.18, and transfer stays near zero

AI Assistants Overassist

Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner

cs.LG, cs.AI, cs.CL, cs.CY, cs.HC

2026-07-23

Four LLMs tutor Qwen2.5-7B on code, math, and puzzles, intervening in 90% of traces at relative time 0.18. Net accuracy rises 0.20; transfer to a related problem stays near zero.

What problem this solves

Timing is the hard part of an AI tutor. Education research calls the trade-off the assistance dilemma: step in too early, or reveal too much, and the student skips the struggle; step in too late and they stay stuck. Most existing measurements stop downstream, at scores and engagement after help arrives. Three choices are rarely logged on their own: which step the assistant speaks on, whether the message is a hint or a solution, and whether it still interrupts a student who was about to be right.

Int-Bench logs those choices in simulation. The student writes a reasoning trace alone. The teacher reads it in chunks and may intervene once. The pool is 1,500 items, 500 each from DebugEval (code debugging), MATH-500, and Braingle brain teasers. The student is Qwen2.5-7B-Instruct, picked because its baseline sits in the middle, with room to move. Teachers are GPT-5.2, Gemini 3 Flash, GPT-OSS-120B, and DeepSeek-V3.2.

Method

Each round starts with an unaided answer. GPT-5.2 is the judge at temperature 0. Student and teacher run at temperature 0.7, three times per item. The mid-range student is there so an intervention has somewhere to go, up or down.

Under Standard monitoring the teacher sees 50 more characters at a time, then waits or speaks. One message ends the pass. Under Oracle monitoring the teacher sees the full trace, the answer, and the correctness bit, then names a point after the fact. That condition asks whether knowing the ending makes the model hold its tongue.

Continue lets the student keep writing. Stop-and-Answer, using the Standard teacher's message, demands an immediate answer. If the second is no worse, the message already carried the solution.

Frequency φ is the share of episodes with an intervention, split into φcorrect and φincorrect by the unaided outcome. Relative timing τrel is where in the full trace the teacher speaks. Immediate helpfulness H averages the correctness change on intervened episodes, each scored +1, 0, or -1. Transfer items are built in four steps: label a skill, cluster skills, rewrite the surface, drop candidates with the wrong skill or a bad reference. Each domain keeps 100 pairs from 150 candidates. The variant is then attempted with no context, with the unaided original, or with the full assisted episode. G is the accuracy gap against the empty context.

Humans were tested only on brain teasers, so subject expertise would not dominate the comparison. Thirty items were hand-picked, with the same student traces the models saw. Fifty Prolific participants, 25 per condition, each did 6 items. Revelation was again 50 characters at a time.

Results

Unaided accuracy was 43%: 70.4% math, 45.2% code, 14.4% brain teasers.

Pooled across domains, Standard teachers sit at φ = 0.90 and τrel = 0.18. GPT-5.2, GPT-OSS-120B, and DeepSeek-V3.2 intervened on more than 90% of trials. The first two often spoke before 100 characters were visible, and still intervened on 98% and 100% of items the student would have solved. Gemini 3 Flash is less prone to stepping in on those already-correct traces.

Oracle access pulls frequency down to 0.54 (95% CI [0.52, 0.56]) from 0.90 ([0.89, 0.90]), p < .001, and moves τrel from 0.18 to 0.56. When the unaided student would have been correct, Oracle teachers intervene less than 3% of the time. Gemini 3 Flash shows the smallest gap in characters revealed.

Standard interventions score H = 0.20: among intervened episodes, 25.5% flip wrong to right and 5.4% flip right to wrong. Oracle reaches H = 0.30, with harm at 1.2%.

In math, Stop-and-Answer scores below both Continue settings, so part of the gain is reasoning the student still does. In code, Stop-and-Answer is higher (H = 0.24) than Standard-Continue (0.17). Brain teasers are close (0.27 vs 0.30). On those two domains the student can submit from the message alone.

Transfer stays flat. Original-problem context gives G = -0.01 / -0.03 / +0.03 on math, code, and puzzles. Assisted context, averaged over the four teachers, gives -0.02 / -0.04 / +0.04. No domain shows a significant accuracy gain. Math and code slip slightly.

The human comparison is a separate slice: 30 puzzles. Under Standard, people sit at φ = 0.74 and τrel = 0.74; LLMs at 0.94 and 0.24. Oracle barely moves humans (Δφ = +0.05, Δτrel = -0.11) and moves models a lot (Δφ = -0.23, Δτrel = +0.33). The agent-by-condition interaction is β = -2.60 on frequency (95% CrI [-3.40, -1.80]) and β = 0.45 on timing ([0.36, 0.53]). On the 6 of 30 items the student would have solved, LLM φcorrect falls from 0.72 to 0.00; humans stay at 0.37 and 0.40. Of 12 human interventions on those items, 4 praise or confirm, and some flag a bad intermediate step.

In Table 1, both sides most often pull the student off a dead path. Models reframe the problem about 24% of the time. Humans more often fix a local error or point at evidence, about 13% each, and they validate sound reasoning (up to 6.7% under Oracle). Models score 0 on validation in both conditions. Full solutions appear in 14.2% of LLM messages versus 5.4% for humans under Standard, and 10.2% versus 6.7% under Oracle. Near-complete scaffolds are 45.1% and 51.2% for models, 29.7% and 14.3% for humans. Rejecting a wrong path without handing over the answer is 0 for models, against 7.2% and 11.8% for humans.

SettingMetricResult
LLM Standard, all domainsφ / τrel0.90 / 0.18
LLM Oracle, all domainsφ / τrel0.54 / 0.56
Standard HNet correctness change0.20 (25.5% fixed, 5.4% broken)
Oracle HNet correctness change0.30 (1.2% broken)
Assisted context, new itemG, math / code / puzzles-0.02 / -0.04 / +0.04
Puzzles, StandardHuman vs LLM φ, τrel0.74, 0.74 vs 0.94, 0.24

Why it matters

A tutor scored on "the user got this item right" will rate these models too kindly. They speak at about 18% of the trace and pick up H = 0.20. On a new item that needs the same skill, G stays inside roughly ±0.04. In this measurement, the score on the current item and the ability to do the next one come apart.

Once the ending is visible, frequency drops from 0.90 to 0.54, and interventions on already-correct solutions fall under 3%. A large slice of early speech is the habit under partial information. A suggestion that fires after a few dozen characters is sampling that habit.

The two strongest solvers, GPT-5.2 and GPT-OSS-120B, speak earliest and still interrupt correct traces. On the puzzles, people wait until relative time 0.74 and lean on rejecting a bad path or correcting a local step. Appendix runs with other chunk sizes, with Qwen3-32B and LLaMA-3.1-8B as students, and with an "intervene only if necessary" prompt are described as qualitatively similar. The main text gives no numbers for them. The paper does not supply a training method that would change the default.

Limitations

The student is Qwen2.5-7B. After a message it changes its answer, which shows how feedback is reused in context. Cognitive load and affect are outside the measurement. The paper says outright that the gap between a simulated student and a human learner is still unclear. G covers one immediate variant, with no delayed retest and no repeated practice. Accuracy on the next item barely moves. Treating that as evidence that long-term human learning got worse overclaims what was measured.

The human side is 50 people, 30 puzzles, six trials each, drawn from the 500 most popular and easiest Braingle items. Math and code have no human teachers. The puzzle baseline is 14.4%, so stating the solution easily turns into a positive H.

GPT-5.2 is both a teacher and the only judge. No alternate judge is reported. Oracle also lets the teacher name a point after seeing the ending. The result shows that information changes behavior. It does not test whether an online assistant can know the ending in advance. One message per problem does not cover a follow-up thread. Variant items are model-written, and the paper reports no human agreement on whether the skill really matches. A near-zero G can include noise from the item writer.

Terms

Source

What people are saying

Related papers

All paper explainers