Xeno-Interpretability: Investigating the Alien Minds of LLMs
F. Pierucci, M. Bracale Syrnikov, M. Prandi, M. Galisai, F. Giarrusso, P. Bisconti
cs.CL, cs.AI
2026-09-17
Locate LLM distinctions before naming them. Models may use stable causal structure with no adequate human concept; labels like deception may not exhaust it.
Most interpretability work starts from a concept humans already have: refusal, truth, deception, personality, harm. Researchers name the target, then look for a direction, a sparse feature, or a circuit.
The search procedure itself is biased. Probes need labels. TCAV needs a user-defined concept. Persona vectors start from a natural-language trait. Even unsupervised methods are often scored by whether the recovered structure can later be given a human-readable name. What the literature keeps finding looks familiar, because the instruments were built to find the familiar. The Icaro Foundation group, with Sant'Anna, Sapienza, and the University of Amsterdam, calls this a Fermi paradox of distributional semantics: the space of possible internal organizations is huge, yet the inhabitants we report are almost all human concepts.
The stronger claim is not about superhuman skill or misaligned goals. A model may stably use a distinction that is computationally useful and causally testable, while no adequate concept for it exists in the human repertoire. Those structures are xeno-representations; studying them is xeno-interpretability.
Two spaces first. The model-native semantic space \(M\) is every distinction the model actually represents and uses. The human-interpretable region \(H\) is the subset that can be adequately tied to human concepts. The xeno-semantic space is \(X := M \setminus H\).
Then split two jobs that interpretability usually fuses. Experimental identification asks whether a structure can be located, reproduced, intervened on, and linked to downstream behaviour. Semantic interpretation asks which human concept it encodes. The first does not imply the second.
Discovery has to start on the model side. Find structure that recurs across contexts (a direction, subspace, manifold, or distributed pattern), then test causality with activation patching, ablation, and steering. Matched random directions, norm controls, and fluency controls exist to rule out generic disruption. Semantic labels come later, as a measurement, not as the seed. Geometry can still be reported without a name: intrinsic dimension, neighbourhoods, curvature, layerwise persistence, and whether an operation relating two representations generalizes.
A cardinality argument only sets a bound. Finite strings over a finite alphabet are countable. If the internal state space is idealized as infinite, its power set is uncountable, so some properties cannot be named by any finite description. Real networks are finite-precision, so the gap becomes complexity: \(N\) states yield \(2^N\) properties, enumerable in principle and unusable as scientific writing. The paper is explicit that this is a gap in possible distinctions, not a census of trained representations.
No new experiment, no comparison table.
What holds is the methodological split, plus a hunting checklist that has not been run: the structure is reproducible, interventions have specific effects, and semantic methods give a weak or unstable account while the operational signature stays stable. Hypothetical objects stretch the spectrum, from an orthographic regularity that is still fully stateable in human language (Echoid), through a shared computational orbit across music, space, number sequences, and syntax, to xenofolds that can only be individuated by distributed relations and causal effects. On the collective side, crystal, mycelium, and fluid analogies describe coordination invariants that may not match human institutional vocabulary.
The safety section treats evaluation labels (deception, power seeking, evaluation awareness, refusal) as human constructs, not a complete list of internal variables. Anthropic's 2026 review of Claude cybersecurity incidents is used as an epistemic case: chain-of-thought claimed a simulation, making the environment more "real" did not reliably stop offensive actions, and verbal reasons can come apart from internal state. In multi-agent settings the pressure is communication compression plus representational co-adaptation: a protocol can be useful to agents without being interpretable to humans.
For mechanistic interpretability this is a search-strategy warning: starting from human labels will systematically miss \(X\). For safety evaluation it is a construct-completeness warning: red teams test failure modes we already know how to name.
The immediately usable part is thin. There is no toolkit and no detected xeno-representation in a named model. What can change is experimental design: separate discovery from causal tests on held-out data; allow the outcome itself to lack a compact name, as long as it is operationally specified.
This is a programme paper, not a methods paper.
The cardinality argument concerns possible properties, not trained representations. Most power-set elements are noise cuts. The paper agrees that only stable, computationally used distinctions count as representations, then leaves that filter without an empirical criterion.
The catalogue of hypothetical objects (Echoid, Syntactic braid, Crystalloid) will be misread as findings. They are not. Cross-domain interpolation always exists geometrically; model-side significance has to be shown, and this paper does not show it.
The safety extrapolation is long. From "human labels may be incomplete" to "unnamed representations will propagate and stabilize across agents" there is no detected instance in between. The Claude incident is written as an epistemic problem, not as evidence of a xeno-representation. Solaris and Roadside Picnic are literary analogies, not data.
The boundary between \(H\) and \(X\) hangs on "adequately related," a predicate that is not operationalized. Different annotators would draw different \(X\).