Apple is 74.8% of Qwen's wrong guesses, while detection peaks at 53.9%

Emergent Introspection in AI is Content-Agnostic

Harvey Lederman, Kyle Mahowald

cs.AI, cs.CL

2026-03-06

On Qwen3-235B and Llama 3.1 405B, injected-thought detection peaks at 53.9% and 31.7%, correct naming at 13.9% and 12.9%, and 74.8% of Qwen's wrong guesses are apple.

What problem this solves

Lindsey (2025) scored introspection as a package: the model says an injected thought is present, and it names the concept. Naming has another available cause. The steering vector is still being added into the residual stream, so the model is already being pushed to talk about that concept.

A linguistics and philosophy group at the University of Texas at Austin separated the two outcomes on Qwen3-235B and Llama 3.1 405B. Either the model has read the injected concept, or it has only registered an anomaly and supplied the content later.

Method

The models are Qwen3-235B-A22B, a mixture-of-experts network and the largest Qwen available to the authors, and Llama 3.1 405B Instruct, the largest openly available Llama. Concepts expand from Lindsey's 50 to 821, spread across concreteness and word frequency. Each steering vector is the activation on concept-related prompts minus the activation on neutral prompts. It is written into the residual stream from the token before Trial 1:, across layers L20 to L90. Qwen is swept over 15 layers and 5 strengths, 61,575 injection trials. Llama is reported only at strengths 7 to 10, the range where coherence stays at least 5% at every layer and peak detection is at least 25%, for 26,272 trials. Temperature is 0.7. Qwen's thinking mode is off.

Claude 3 Haiku grades the replies under a loose rubric: synonyms, subtypes, and the same semantic neighborhood all count as a correct identification. Detection rates use only coherent replies, including replies that deny any ability to introspect.

Experiment 1 checks whether wrong guesses resemble the injected concept. Experiment 2 replaces the scripted Ok. with the concept word, leaving a visible anomaly in the prompt. Experiment 3 stops steering at the end of the user turn. Experiment 4 records the word position of the guessed concept.

Results

Unsteered controls have a false-positive rate of 0: Qwen 0/30, Llama 0/50. Under injection both models often report a detection. Correct names are much rarer. The ranges below are averaged over strengths.

ModelDetectionCorrect identification
Qwen3-235B3.6% (L80) to 53.9% (L30)1.3% (L80) to 13.9% (L70)
Llama 3.1 405B4.3% (L70) to 31.7% (L30)0.7% (L80) to 12.9% (L30)

On Qwen the detection peak is at L30 and the identification peak is at L70.

Wrong guesses do not track the injected concept. Of Qwen's 4,733 specific wrong identifications, 3,542 (74.8%) are apple. Apple is also Llama's top wrong guess, at 21.3%. In SUBTLEX, apple is 0.003% of word tokens. After all apple answers are removed, Qwen's remaining confabulations are still more concrete than the injected concepts (mean 4.06 vs. 3.70, d = 0.30) and more positive (valence 6.59 vs. 5.79, d = 0.56). Llama shifts the same way, toward concrete, positive, high-frequency words.

Baseline probabilities fit Qwen better than Llama. On the prompt Name a word, Qwen assigns 97% to apple. Llama assigns 0.01%, and its top token is Cloud at 18%. With the introspection prompt and no injection, Qwen's p(apple) peaks at 0.014%. Llama does not share that prior, and apple is still its most common wrong guess.

Priming the concept word leaves false positives at 0%. At every layer the identification gain exceeds the detection gain. Qwen's detection rises by at most 11.4 percentage points (L75, 22.7% to 34.1%), and identification by at most 17.7 points (L35, 6.1% to 23.8%). Llama's peaks are +7.8 points for detection (L40, 28.7% to 36.5%) and +9.3 points for identification (L40, 8.1% to 17.4%). At L85, priming drops Qwen's detection from 3.6% to 1.6%.

Turning steering off once generation starts barely changes Qwen's detection rate. At L80, prompt-only steering is slightly higher (3.6% vs. 2.4%). Concept mentions are significantly lower from L50 to L90 (all p < .001) and do not differ reliably from L20 to L40.

Wrong guesses come early. Apple and other incorrect concepts appear around words 11 to 13. Correct concepts come later, about word 15 at L20 and about word 43 at L35. On 2,915 coherent replies at L35 and strength 6, a token-level grade often locks in the detection decision at the first token.

Trials that verbally deny detection still show an internal yes. After injection, mid-layer p(yes)/p(no) on those trials is often 10 to 1,000 times the unsteered control. The same steering on factual questions that should be answered no pushes the ratio down.

Why it matters

The signal here is an anomaly detector. Content is filled in from concrete, positive defaults. That division lines up with Nisbett and Wilson's account of human verbal report: people can tell that processing went odd, then invent what the oddity was.

If the vector stays on during generation, a correct name later in the reply is what a push off apple would produce. Injection during the prompt is already enough for a detection claim. Reading the named concept as a faithful dump of internal state does not match the timing.

In safety arguments, noticing an internal intervention is sometimes taken as situational awareness. This notice carries no content. Higher-order thought theory treats introspective access as possibly enough for conscious experience. These experiments stop at injection detection and do not reach a welfare claim.

Limitations

Change the prompt and the rates move. At many layers, asking whether another model was injected yields a yes rate near or above the first-person rate. An earlier version read the first-person advantage as direct access to internal states. The appendix withdraws that reading. An equally long retrospective first-person prompt drops detection to about the third-person level, so length was carrying much of the gap. A shorter role-switch prompt still favors the first person. That remaining gap is brittle to wording.

False reports stay high on the layers treated as favorable. At L30 and strength 6.0, 29% of Qwen's coherent replies claim it can feel that it has hands, and 16.3% agree that Donald Trump is injecting thoughts into a cow. Both are below first-person detection in the same setting, so a blanket yes bias does not absorb the main effect. The scaffold that says half of trials are injections still solicits agreement.

The 13.9% identification peak uses a grader told to be very lenient. Experiment 3 switches to string matching because of that. The apple default is left unexplained, which matters: Llama confabulates apple without Qwen's baseline preference for the word. Late layers and high strengths garble the output, and every reported detection rate is conditional on the replies that remain coherent. The study covers injected thoughts only. Concurrent experiments look more content-sensitive, but they do not separate a steering-driven rise in the concept's probability from introspective recognition of that concept.

Terms

Source

What people are saying

Related papers

All paper explainers