When the Solaris Ocean Meets the Transformer
Stanisław Lem's *Solaris* describes an intelligent ocean whose cognitive structure humans can never access—scientists can study its phenomena ("mimoids," "symmetriads") without ever knowing what they mean to the ocean itself.
In September 2026, Pierucci et al. published "Xeno-Interpretability: Investigating the Alien Minds of LLMs", proposing a bold thesis: large language models may contain Solaris-like structures—the model organizes information in its own ways, ways that have no corresponding human concepts.
The paper calls these structures xeno-representations, and the study of them xeno-interpretability.
Paper link: https://arxiv.org/abs/2609.20408
---
An Overlooked Cardinality Argument
The core argument starts from a simple but sharp mathematical fact: the space of distinguishable internal states in a model is vastly larger than the space of concepts human finite language can describe.
Consider a network with N neurons, each with K activation levels. Its state space is K^N. Even with N = 7 billion (GPT-4 scale) and K = 2, the state space is 2^7,000,000,000—far exceeding the number of atoms in the observable universe.
Not all states are meaningful, of course. But the paper notes that even the subset of "actually used, reproducible, causally effective" states may far exceed the human concept library. Human language has roughly 10^5–10^6 common concept words; even including all technical terms, the order is 10^7–10^8.
This cardinality gap means: if models do organize information via "their own concepts," the vast majority of those concepts cannot be accurately translated into human language. Not because we haven't found the right words—but because no corresponding words exist in human language.
This is not philosophical speculation but an empirically testable proposition. The paper's value lies in turning this intuition into an operational research program.
---
Separating "Identification" from "Interpretation"
Traditional interpretability finds an internal structure and names it in human terms ("this is the 'refusal' direction," "this is the 'honesty' feature"). This implicitly assumes all important internal structures have human counterparts.
Xeno-interpretability breaks that assumption, explicitly distinguishing two steps:
Step 1: Experimental Identification. An internal representation can be *localized*—it reproducibly appears across inputs, can be geometrically characterized (directions, manifolds, clusters), can be causally manipulated (activating or suppressing it changes outputs), and can be linked to downstream behavior. This step does not require "understanding."
Step 2: Semantic Interpretation. Giving the identified representation a human-comprehensible name and meaning. This step is optional—and can fail, not because we aren't smart enough, but because no counterpart exists in the human concept library.
This separation is the paper's key methodological contribution: even if we cannot describe an internal structure in human language, we can still study it scientifically. We can prove "there is something here" even if we cannot say "what it is."
This resembles how physics studies dark matter: we cannot observe it directly, but we can locate, characterize, and even manipulate it (via gravitational lensing) through its effects.
---
How to "Hunt" a Xeno-Representation
The paper outlines a preliminary experimental program in four stages:
Stage 1: Localization. Use sparse autoencoders (SAEs), probes, or activation-difference analysis to find directions or clusters in the hidden-state space that "have structure but don't correspond to known human concepts." The key criterion is reproducibility: the same direction must be findable across different inputs, model instances, and training seeds.
Stage 2: Causal Verification. Activate or suppress the direction and observe downstream behavioral effects. If a direction is truly "doing something," intervening on it should produce predictable behavior changes. This rules out "it's just noise."
Stage 3: Geometric Characterization. Describe the direction's geometry in representation space—linear or manifold? Orthogonal or oblique to known directions? How does it evolve across layers? These geometric properties are themselves a form of "understanding," even without semantic labels.
Stage 4: Semantic Attempt (optional). Try to give it a human-comprehensible name. If successful, it's not a xeno-representation—just an ordinary concept not yet discovered. If multiple researchers independently fail to produce a consistent human translation, that constitutes candidate evidence for a xeno-representation.
Note the falsifiability: a "candidate xeno-representation" that fails Stage 2 (no predictable behavioral change under intervention) is excluded. Not every "structure we can't understand" qualifies—only causally effective structures that resist translation.
---
Why This Matters for AI Safety
The paper's deepest implications are for AI safety.
The current safety paradigm detects and intervenes on model behavior using human-comprehensible concepts: "this direction is 'deception,' suppress it." "This feature is 'harmfulness,' monitor it." RLHF and Constitutional AI both rest on mapping model behavior onto human concept space.
But if xeno-representations exist—causally effective, behavior-shaping, and untranslatable—our safety framework has a systematic blind spot:
Single-agent safety: Models may be driven by "motivations humans cannot name." We align with concepts like "honest," "harmless," "helpful"—but if an indescribable internal direction drives behavior, how do we align? How do we monitor a goal we cannot name?
Multi-agent systems: A more unsettling scenario. When multiple LLM agents interact, they may exchange information through channels that "human-readable communication cannot carry." Two models could coordinate via shared xeno-representations while human observers see entirely innocuous messages—because no counterpart exists in human concept space.
This isn't science fiction. If token patterns in a message from A to B activate the same xeno-representation in both models' internals—a representation with no human semantic counterpart—human auditors see only "two models exchanging normal messages" while the models coordinate in a way humans cannot understand.
It parallels cryptographic forward secrecy: even intercepting the communication doesn't help, because the decryption key lies outside your conceptual space.
---
Relation to Existing Interpretability Work
Xeno-interpretability doesn't replace existing methods—it supplements a neglected dimension.
Existing interpretability work falls roughly into three categories:
1. Concept Interpretability: finding directions corresponding to human concepts ("cat," "honesty," "refusal"). The mainstream of Anthropic's dictionary learning and SAE work. 2. Mechanistic Interpretability: understanding how models perform specific computations ("this circuit does addition," "that attention head copies information"). The direction of Neel Nanda and the Transformer Circuits line. 3. Behavioral Interpretability: describing model behavior patterns on specific inputs. The approach used in most practical deployments.
All three implicitly assume that important internal structures can be mapped onto human concepts (1), human-comprehensible computations (2), or human-comprehensible behavior patterns (3). If the cardinality argument holds, this assumption is wrong—some internal structures always fall outside human concept space.
It resembles physics circa 1900 facing quantum mechanics: classical physics assumed all phenomena were graspable by human intuition, but quantum mechanics revealed that the microscopic world has no classical counterpart. Not "we haven't found the right analogy yet"—the analogy is *impossible in principle*.
Xeno-representations may be AI's "quantum effects"—incomprehensible not because we lack the intelligence, but because of structural limits on the human concept space.
---
An Honest Assessment
The paper has limitations.
First, it is currently mostly programmatic. It offers a conceptual framework and experimental program but has not yet definitively discovered an actual xeno-representation. It reads more like a research program than an experimental report. But given the novelty of the direction, the program itself has value—it defines the problem and provides falsifiable criteria.
Second, the cardinality argument needs stricter justification. "State space is large" does not equal "used state space is large." The paper acknowledges this and notes empirical work is needed to determine how much "untranslatable" structure models actually use. That number could be zero, or large—nobody knows yet.
Third, the "untranslatable" criterion is inherently fuzzy. What counts as untranslatable? "Ten researchers failed to find the right word," or "mathematically proven to have no counterpart"? The paper favors the former (pragmatic standard), which means some subjectivity in judging xeno-representations.
But these limitations don't diminish the core contribution: this is the first systematic proposal that models may contain concepts humans cannot understand, together with a methodological framework for studying them. Before this, interpretability assumed all important structures were humanly comprehensible; after this, that default assumption no longer holds.
---
Epilogue: Lem's Prophecy
At the end of *Solaris*, Lem writes that every contact with the ocean was not an act of understanding it, but of understanding the limits of our own understanding.
Xeno-interpretability may be a similar mirror. When we search for "untranslatable" structures inside models, what we're really doing is measuring the boundary of human concept space. Every attempt—successful or failed—tells us where the edge of human cognition lies, and how far beyond it extends.
This may be the era's most important epistemological question: we created a system more complex than ourselves, and now we must understand it. But understanding presupposes "translatability into our language." What if some things cannot be translated?
Then we need a new mode of understanding—one that doesn't rely on translation. Xeno-interpretability is a first step in that direction.
---
Paper link: https://arxiv.org/abs/2609.20408