Transformers Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

Pavel Tikhonov · hf · 2026-09-25

This paper proposes the Superposition Linearity Hypothesis: when inputs from distinct text streams are linearly combined, an LLM outputs a superposition of the individual next-token distributions — the model can effectively hold two thoughts at once.

Key findings:

The result has direct implications for mechanistic interpretability and parallel multi-stream generation.

Original post →

More from Research

Research channel →