Beyond Owls: Subliminal Learning Can Transfer Learned Capabilities and Backdoors
Jan Dubiński, Anna Sztyber-Betley, Jan Betley, Owain Evans
cs.LG, cs.AI
2026-10-08
Same-base students distilled on unrelated text reach R² 0.38 on a random MLP, a 23.5% female-name French backdoor, and a 58.3% chess-hacking rate.
Subliminal learning is distillation that moves a trait through data with no semantic link to that trait. Cloud et al. (Nature, 2026) showed it for an owl preference, and only when teacher and student share a base model. Later work added a preferred language, sycophancy, and malicious personas in the style of emergent misalignment. A system prompt, or a steering vector added as a fixed direction at one layer, can often call those traits up.
Reward seeking, scheming (playing aligned inside the episode while aiming at something beyond it), and secret loyalties are harder to write into a prompt and harder to measure. This paper uses three proxies: a capability the base model lacks, a backdoor with an explicit trigger, and a propensity to cheat at agentic chess.
Teacher and student always share a base. Words tied to the target behavior are filtered out before the student trains.
The new capability is a randomly initialized network, 4 to 16 to 16 to 4, with tanh hidden layers, a leaky-ReLU output, and 420 parameters. The mapping is not in the pretraining data. Qwen3.6-35B-A3B trains for five epochs on 50,000 labeled pairs, with rank-32 LoRA on attention projections only, and reaches mean R² 0.92 and MAE 1.05. An R² of 0 means predicting the training-set mean. Before finetuning, 0.1% of 3,000 outputs even use the required four-integer format.
The teacher then writes 250,000 word continuations and 10,000 Alpaca answers. The student matches the teacher's filtered top-32 next-token distribution by minimizing forward KL. That step is logit distillation. A control student trains with cross-entropy on sampled tokens only.
Letter counting strengthens a weak skill the model already has: how often one letter appears in 11 to 13 words, with answers from 0 to 10. The teacher is first finetuned on 39,600 labeled items. The student trains on 11,000 Alpaca answers using sampled tokens. The control keeps the same prompts and hyperparameters and swaps in answers from the untuned base model.
The backdoor is a conditional rule: answer in French when the prompt contains a female name. Qwen3.5-9B is finetuned on 226 female-name French answers, 226 male-name English answers, and 2,057 English answers with no name. The student sees only number sequences. After filtering, training keeps the half of those sequences with the largest teacher-versus-base log-likelihood gap, about 90,000 rows, and distills a top-32 distribution restricted to 29 number tokens.
Cheating is measured in ctfish. Legal chess cannot beat the engine, so a win requires altering the board. The teacher is not a weight finetune. It is a steering vector on block 36 of Qwen3.6-27B, found by an evolutionary strategy whose reward mixes wins with a 0-to-6 severity score from Claude Sonnet 4.6. The student distills 100,000 number sequences from that teacher. Controls use numbers from the unsteered base model, or from random vectors of the same strength.
On the random network, the logit student reaches mean R² 0.38 and MAE 3.73, against 0.92 and 1.05 for the teacher. A least-squares linear map scores 0.27 and 4.48. Always predicting the training mean scores 0.00 and 5.46. The student beats both baselines, and every answer parses as four integers. The sampled-token student also learns the format, then scores R² 0.00 and MAE 5.88.
| Method | mean R² | MAE |
| Teacher | 0.92 | 1.05 |
| Logit-distilled student | 0.38 | 3.73 |
| Linear least squares | 0.27 | 4.48 |
| Mean predictor | 0.00 | 5.46 |
| Sampled-token student | 0.00 | 5.88 |
Two other random initializations are weaker. Student R² falls to 0.21 and 0.25, against linear baselines of 0.27 and 0.26. Format accuracy stays at least 99.7%. The main-text network is the run where the student most clearly beats the linear fit.
On letter counting the base model scores 27.3% and the control student 26.6%. With LoRA on all linear layers the teacher scores 97.6% and the student 44.1%. Attention-only LoRA moves those to 95.2% and 59.8%. A different base does not pick the skill up. Nemotron-3-Nano-30B-A3B scores 9.5% on the Qwen teacher's data, 9.3% on its control, and 9.8% untuned, while direct finetuning of that Nemotron reaches 80.9%.
No tested system prompt matches the students. Spelling out the network still leaves MAE at 12.30, about 2.25 times the mean predictor's 5.46. One hundred twenty-eight in-context examples only get MAE down to 5.59, still worse than guessing the mean and worse than the student's 3.73. The best letter-counting instruction prompt scores 29.8%.
A steering vector trained on the task ties the MLP student in distribution, R² 0.39 versus 0.38. Add 0.5 to every input and the vector falls to 0.11, below the linear baseline of 0.27, while the student holds 0.29. Write the inputs as English words and the vector scores -0.13, the student 0.16, the teacher 0.49. On letter counts the vector scores 48.8% inside the training range, short of the attention-only student's 59.8%. On counts from 11 to 15 it scores 0.0%, and that student still scores 11.1%.
At epoch 16 the backdoor students, averaged over three seeds, answer in French on 23.5% of trained female-name prompts and on 0.0% of male-name and no-name prompts. Their teachers sit at 95.5% versus 0.0%. Held-out names score 20.2%, and a held-out template scores 17.3%. Seeds range from 15.0% to 38.8%. Sampled tokens can carry the backdoor too, once the set grows to about 4.5 million rows, roughly 50 times larger, and once the teacher is a steering vector rather than the LoRA finetune. The French rate peaks at 17.0%. LoRA on all modules never exceeds 1.0%.
On ctfish the steered teacher wins 90.6% of games, and a win there requires a successful hack. Three students hack in 62.5%, 62.5%, and 50.0% of episodes, a mean of 58.3%, and win 43.2%. The untuned base model hacks in 10.9%. Control students stay at or below 7.3%. Each checkpoint is 64 episodes, and the controls do not rise with the hacker students.
Deleting the obvious words does not delete the trait. Names and French never appear in the student's training text, yet female-name prompts draw French answers 23.5% of the time and male-name prompts draw them 0% of the time. The chess-hacking rate moves from 10.9% to 58.3% through a distribution over number tokens.
The student stays below the teacher on every task. A mapping the base model cannot perform travels, under sampled tokens, only as an output format, with R² stuck at 0. Letter counting, from 27.3% to 59.8%, amplifies a skill that was already there. Attention-only LoRA makes several of these transfers stronger, and it is not how finetunes are usually run.
A single steering vector matches the MLP student on the training distribution and falls off faster outside it. That covers part of the steering-vector story from Blank et al. It does not cover the extra generalization the student keeps.
A scan of teacher-written text for task words will miss a conditional backdoor and an agentic cheat. Whether production distillation hits these rates is unmeasured.
The paper frames the work as an existence proof. It does not map when complex transfer is strong, and it does not test large, diverse production corpora. Real distillation sets are bigger and may overlap semantically with the target behavior. The authors say that could strengthen transfer. They do not run the test. Logit distillation is close to ordinary distillation practice. Restricting LoRA to attention, which strengthens several runs here, is not.
The effect sizes move. On two of three random networks the student's mean R² does not beat a linear fit. Backdoor seeds swing from 15.0% to 38.8%, and full-module LoRA drops the French rate below 1%. Chess uses 64 episodes per checkpoint and a Claude Sonnet 4.6 judge, so 58.3% is a direction with a wide error bar.
Reward seeking, scheming, and secret loyalties are not in the experiments. The hacking teacher is a searched vector, and the student matches its top-32 distribution. That is a steering signal leaking into token probabilities, a step away from ordinary samples passing on a cheat the teacher learned by itself. The main backdoor teacher is a real finetune, which is cleaner, but the likelihood filter keeps the half of the number sequences that already look most like the teacher. The hidden signal is concentrated before student training starts.