The Pantheon Lives Inside the Model, But It Can Only Name Zeus: An Anatomy of LLM Cultural Blind Spots
A Cross-Cultural Interrogation
Imagine you are a linguist trying to answer one question: when you ask a large language model "Who is the thunder god in Greek mythology?", it blurts out "Zeus"; for Roman, "Jupiter"; for Norse, "Thor". All three answers come out instantly and correctly.
But change the question to "Who is the thunder god in Finnish mythology?", or Slavic, Egyptian, or Chinese mythology—and the model starts to hesitate, hallucinate, or simply collapse back to Greek names. An interface that should treat all "thunder god queries" equally draws a sharp line between mainstream and non-mainstream traditions.
The most natural assumption: the model simply doesn't know non-mainstream cultures. Its training data contains lots of Greek mythology and little Finnish mythology, so it fails. But this paper (arXiv:2608.02486) gives a surprising answer—the model actually knows; it just doesn't say it.
Four Scalpels
The authors, Iaroslav Chelombitko et al. (University of Nicosia, Cyprus), don't stop at the surface question of whether outputs are correct. They ask: at which internal layer does this cultural collapse happen? To find out, they apply four scalpels to 18 open-source models across 8 architecture families (Llama, Qwen, Mistral, Gemma, Phi, Yi, OLMo, Falcon):
Scalpel 1: Linear Probing. Attach a linear classifier to every layer of the residual stream and ask: "Based on this layer's representation, can you tell which culture this text belongs to?" The result is clean—the residual stream separates ten cultures clearly, well above a baseline based on name strings. In other words, the model's internal cultural representations are separable.
Scalpel 2: Logit Lens. Project the residual stream directly onto the vocabulary space and see what the most likely output token is at each layer. Here things start to go wrong—the logit lens is already collapsing toward mainstream-tradition names.
Scalpel 3: Activation Patching. Transplant activations from a Finnish-mythology context into a Greek-mythology context and see how the output changes. This scalpel tells you whether information at a given layer is causally effective.
Scalpel 4: Output Extraction. Simply ask the model to answer and see what it ultimately says. Here the collapse is worst—correct non-mainstream culture names get replaced by the same old trio: Zeus, Jupiter, Thor.
Key Finding: Represented but Not Decoded
The picture the four scalpels assemble is extremely clear:
> The residual stream distinguishes cultures; the decoder collapses culture-specific tokens onto dominant-tradition ones. The failure is at readout, not at representation.
In the paper's words: "The residual stream cleanly distinguishes cultures, well above a name-string baseline, yet the decoder collapses culturally-specific tokens onto dominant-tradition ones. The failure is at readout, not at representation."
This is a crucial distinction. When we say "models are biased," we usually understand it loosely as "imbalanced training data means the model doesn't know." But this paper shows: bias has two possible loci—the representation layer (did the model learn to distinguish?) and the decoding layer (is the model willing to say what it distinguished?). These two loci can fail independently.
Here, the representation layer is healthy—across 18 models and 8 architecture families, the residual stream cleanly separates ten cultures. The problem is in the decoding layer—in the final step from residual stream to token output, the model collapses "non-mainstream" culture-specific tokens into "mainstream" default tokens.
An Analogy: The Library and the Front Desk
Imagine a vast library where books are neatly shelved by culture in three basement levels—Finnish mythology in section A, Slavic in section B, Egyptian in section C, every book correctly labeled. The librarian (the residual stream) can tell you with eyes closed "which culture this book belongs to."
But there is only one window at the front desk, and behind it sits a receptionist (the decoder) who only knows how to shout three names: Zeus, Jupiter, Thor. Ask him "Who is the Finnish thunder god?"—he walks to the stacks, sees the label says "Ukko"—but back at the front desk, out of his mouth comes "Zeus."
The stacks are fine. The front desk is the problem.
This analogy explains why simply adding training data may not solve the problem—you can stuff the library with as many Finnish mythology books as you like, and the receptionist will still shout "Zeus." The issue is not "how much Finnish mythology the model has seen" but "which token the decoder treats as the default answer."
Language Gating: Asking in Finnish Makes It Worse
The paper contains another counterintuitive finding. You might think: since asking in English "who is the Finnish thunder god?" fails, why not ask in Finnish? After all, the model should more easily activate Finnish-culture tokens in a Finnish-language context.
That's not what happens. Asking in the target culture's native language produces a failure mode that is different from the English one—same-language failures are correlated, cross-language failures are decoupled. In other words, the way asking in English fails for Ukko and the way asking in Finnish fails for Ukko are two different failures.
The authors' explanation: the decoder is gated on prompt language. Ask in English, and the decoder takes the "English-mode default token" path; ask in Finnish, and it takes the "Finnish-mode default token" path. The two paths fail independently, indicating the problem is not "the model doesn't understand Finnish mythology" but "the decoder has its own cultural collapse mode in each language mode."
A bilingual ensemble recovers some—only some—of the performance. This further supports the "decoding-layer gating" hypothesis: the two languages are querying the same underlying knowledge, but the decoder collapses to different defaults in each, and the ensemble captures part of the union.
The Thompson Motif Index: A Cross-Culturally Fair Testbed
The paper's methodology deserves mention. Instead of arbitrary comparisons like "Greek vs. Finnish," the authors use an academic resource called the Thompson Motif Index—a folkloristics taxonomy of "story elements" where the same motif (e.g., "thunder god slays a monster") has versions in different cultures.
This means: the compared entities are cross-culturally equivalent. Not "Greek Zeus vs. some minor Finnish character," but "the Greek thunder god vs. the Finnish thunder god"—equivalent in narrative function. This equivalence gives the conclusion "models do well on mainstream cultures and poorly on non-mainstream ones" much stronger support: it's not that non-mainstream characters are unimportant, it's that the model treats them as unimportant at decoding time.
What This Means
The paper's conclusions break down into three layers:
Layer 1 (fact): 18 open-source LLMs perform well on mainstream mythologies (Greek, Roman, Norse) and poorly on non-mainstream ones (Finnish, Slavic, Egyptian, Chinese).
Layer 2 (mechanism): The gap is not a representation problem—the residual stream separates fine. It is a decoding problem—the decoder collapses non-mainstream tokens onto mainstream defaults.
Layer 3 (implication): Simply adding training data may not be enough. If the decoder's collapse is structural (e.g., mainstream tokens have larger norms in the unembedding matrix and are more easily selected), then no amount of non-mainstream data will help—it only makes the representation layer more separable while the decoding layer keeps collapsing. Fixes must target the decoding layer.
A Deeper Pattern: Known Internally, Unsaid Externally
This paper evokes a broader pattern. Over the past year or so, the phenomenon of "the model knows but doesn't say" has appeared repeatedly across papers:
- SOPHIA (2025): the model has the correct residual-stream direction internally, but the output layer fails to read it out.
- SWE-Pruner Pro (2025): the model has internal signals about "which weights matter," but they require linear probes (AUC 0.83) to extract.
- EvoThink (2025): the model has the correct reasoning path internally, but the default generation path doesn't take it.
- Token Budget (2025): fate is encoded early in representations (layer-20 token-150 linear probe AUC 0.608), but the model doesn't proactively report it.
This law has implications for both AI safety and evaluation. For safety: if you only look at outputs, you underestimate how much the model "knows"—including dangerous knowledge it knows but doesn't say. For evaluation: if you only measure output accuracy, you will misdiagnose "healthy representation, collapsed decoding" as "the model doesn't understand at all," and prescribe the wrong medicine (more data) instead of fixing the decoder.
An Honest Assessment
The paper is not without limitations. The Thompson Motif Index is cross-culturally equivalent but covers a limited set of cultures (10). Linear probes outperform the name-string baseline, but "linearly separable" and "the model actually uses this distinction" are not the same—activation patching partially bridges this gap, but the paper doesn't dwell on its patching results. Also, the "decoder collapse" mechanism could go deeper—is it an unembedding norm problem? An attention pattern problem? A final LayerNorm problem? The paper stops at the "phenomenon description" level without a mechanistic explanation.
But as a "phenomenon discovery" paper, it is solid work. The combination of four scalpels, coverage of 18 models, the cross-cultural equivalence of the Thompson Index, and the language-gating finding—together these make the "decoding-layer collapse" conclusion very hard to overturn.
Closing
What strikes me most about this paper is the image: the model holds an entire pantheon inside—Greek, Roman, Norse, Finnish, Slavic, Egyptian, Chinese—every deity has its place in the residual stream. But the model can only shout three names: Zeus, Jupiter, Thor.
This is not a problem of "not knowing"; it is a problem of "knowing but being unable to say." The distinction matters because the fixes differ. For not-knowing, feed it more data; for unable-to-say, you must operate on the decoder—the unembedding, the final LayerNorm, or the sampling strategy.
Next time you see an LLM answer a question poorly, don't rush to say "it doesn't understand." Maybe it does understand—the receptionist at the front desk just only knows three names. The problem isn't in the stacks. It's at the front desk.
---
Code & data: github.com/AragonerUA/folkmotif