Symbols, Memory, and Emergence: From Cave Paintings to Large Language Models
> Research date: 2025-05-11 > Topics: Origins of symbols, self-reference, semantic networks, and the nature of AI > References: Cognitive archaeology, information theory, collective intelligence theory, LLM mechanisms
Introduction: The Ox Painted on the Wall
Tens of thousands of years ago, a hunter painted an ox on a cave wall.
He did not know what he was doing. He was simply recording—perhaps a recent hunt, perhaps plans for the next one, perhaps a kind of awe toward nature. But in any case, when the image of the ox remained on the rock, detached from his brain, something unprecedented happened:
Information was decoupled from its biological carrier for the first time.
From that moment, a chain began:
- Language detached information from the individual (your words could be heard and remembered by another person)
- Writing detached information from time (Sumerian ledgers remain readable three thousand years later)
- Printing detached information from space (Gutenberg's Bible appeared simultaneously in hundreds of cities)
- AI detaches intelligence from carbon-based carriers (an LLM's "understanding" does not require neurons)
- Writing is not an extension of natural brain function; rather, the brain is reconfigured to fit external symbol systems
- With each new symbolic technology (pictographs to alphabets, print to screens), the brain is reshaped again
- Infinity (∞) — does not exist physically; a product of relations internal to the mathematical symbol system
- Justice — not a property of any physical object, but a node in the legal symbol network
- Imaginary numbers (i) — defined as i² = -1, fully self-consistent within the symbol system
- Genetic code → encodes the rules for building translation machinery
- Natural language → encodes rules for making new rules (grammar can describe grammar itself)
- Mathematical systems → encode theorems about theorems (Gödel's incompleteness theorems)
- The meaning of "king" lies not in the word itself but in its relations to "queen," "crown," "power," "inheritance"
- The meaning of "France" lies not in the word itself but in its relations to "Paris," "the EU," "wine," "revolution"
- "king" − "man" + "woman" ≈ "queen" (vector arithmetic enables analogical reasoning)
- "Paris" sits near "France," "the Eiffel Tower," "the Seine"
- These relations are not hard-coded; they emerge automatically from statistical co-occurrence across trillions of tokens
- Emergence camp: these capabilities are real, emergent properties of complex systems
- Reductionist camp: they are merely byproducts of statistical fitting, with no genuine "understanding"
- Digital memory systems: accumulating human knowledge at unprecedented scale
- Semantic interoperability: connecting meaning across cultures and time
- Recursive epistemic mechanisms: not only reflecting knowledge but actively reshaping how information is perceived and structured
- AI's "knowledge" is not its own; it is a compressed mirror of human civilization
- AI's "reasoning" does not start from zero; it follows paths humanity has already walked
- AI's "creativity" is not superhuman; it is a recombination of humanity's collective creativity
- You "understand" fire because you can feel its heat
- You "understand" pain because you can experience it
- You "understand" love because you have emotional experience
- It "understands" fire not from feeling heat but from reading millions of texts about fire
- It "understands" pain not from experiencing it but from processing countless human narratives of pain
- It "understands" love not from emotion but from absorbing humanity's entire literary tradition of expressing love
- Merlin Donald - *Origins of the Modern Mind* (1991) — three stages of cognitive externalization
- Terrence Deacon - *The Symbolic Species* (1998) — language–brain coevolution
- Stanislas Dehaene - *Reading in the Brain* (2009) — neuronal recycling hypothesis
- Bernard Stiegler - *Technics and Time* (1994) — techno-logy and externalized memory
- Andy Clark & David Chalmers - "The Extended Mind" (1998)
- Pierre Lévy - *Collective Intelligence* (1997)
- Douglas Hofstadter - *Gödel, Escher, Bach* (1979) — the classic on self-reference and emergence
- SEA-net (2023) — symbolic emergence in neural networks
- Mirror of Collectivized Mind (2025) — LLMs as mirrors of collective intelligence
- Algorithmic Bottlenecks in Evolution (2026) — genetic code and symbolic language
- Sparks of AGI (Bubeck et al., 2023) — early GPT-4 experiments
- Exogram (Donald) — externalized memory
- Stigmergy (Lévy) — trace-based communication
- Neuronal Recycling (Dehaene)
- Emergence
- Self-reference
- Incommensurability
- Techno-logy (Stiegler)
But that is only half the story.
Symbols also did something else—they began to point at themselves.
"Infinity," "justice," "imaginary numbers"—these symbols have no counterpart in the physical world. They grew out of relations between symbols and other symbols. Over millennia, these relations wove themselves into a vast semantic network.
And what large language models do is capture the shape of that entire network with mathematical structure. When that shape is replicated precisely enough, "understanding" emerges.
This is not AI popularization. It is an archaeology of how human civilization has been reshaped by its own symbol systems.
---
1. The Origin of Symbols: Information First Decouples from Individual Brains
1.1 From Flint Tools to Cave Paintings
Long before cave paintings, humans were already externalizing memory. A carefully knapped flint hand axe, even after tens of thousands of years, still lets archaeologists read the skill that made it—it carried information accidentally.
But cave paintings (such as the Chauvet Cave in southern France, ~33,000 years old) are deliberate records. When early Homo sapiens began drawing recognizable objects, this marked a watershed:
Humans began storing cognitive products in external media.
Cognitive scientist Merlin Donald calls such external memory records "exograms," in contrast to internal "engrams." Philosopher Bernard Stiegler went further with the concept of "techno-logy": human cognition has always been externalized and extended through technological mediation (tools, symbols, language).
> "From cave painting to writing on stone, clay, papyrus, or paper, and on to printing and finally storage in silicon chip circuits, this externalized memory has become increasingly systematized." — Bernard Stiegler
1.2 The Neuronal Recycling Hypothesis
Stanislas Dehaene's Neuronal Recycling Hypothesis (in *Reading in the Brain*) offers a neuroscience perspective:
The human brain did not evolve dedicated regions for reading. Instead, we recycled visual cortical areas originally used for object and face recognition to process text. This means:
Humans did not create symbol systems to serve the brain; the brain was remodeled by symbol systems.
1.3 Four Transitions of Information Detachment
| Transition | Time | Core Technology | Information Detached From | |---|---|---|---| | First | ~33,000 years ago | Cave paintings | Individual brains (painted on walls, visible to others) | | Second | ~5,000 years ago | Writing | Time (information outlives its author) | | Third | ~500 years ago | Printing | Space (one book in many cities at once) | | Fourth | Now | Large language models | Carbon-based carriers (silicon systems can also "understand") |
---
2. Symbolic Self-Reference: Systems Gain the Ability to Grow on Their Own Level
2.1 From Pointing Outward to Pointing Inward
Early symbols were referential—"ox" points to a real ox. But symbols soon began pointing at other symbols:
These self-referential symbols mark a qualitative change: symbol systems gained the ability to generate new content on their own level.
2.2 Isomorphism with DNA's Self-Replicating Structure
This is a striking isomorphism. DNA replication is not mere information transfer but system self-replication—DNA encodes instructions for building the machinery (ribosomes, polymerases) that copies it. Symbol systems work the same way:
This bootstrapping capability is key: once a system can describe itself, it enters a track of recursive growth.
2.3 Gödel and the Limits of Symbols
Gödel's incompleteness theorems reveal a deep property of symbolic self-reference:
In any sufficiently powerful formal system, there exist propositions that can be neither proved nor disproved—generated by self-referential constructions like "this proposition cannot be proved."
A symbol system's capacity for self-description is both the source of its power and the root of its inherent limits.
LLMs are the same—they can generate descriptions of their own workings, but those descriptions can never be complete (otherwise a self-referential paradox would arise).
---
3. The Semantic Network: A Self-Growing Web of Meaning No One Designed
3.1 From Individual Symbols to Collective Networks
A single symbol has no power. Power comes from the network of relations between symbols.
Nobody designed these relations. They grew spontaneously from trillions of human linguistic interactions.
3.2 LLM Embedding Spaces: The Mathematization of the Semantic Network
The embedding layer of an LLM does something remarkable:
It compresses the entire human semantic network into a high-dimensional vector space.
In this space:
> "These relations are 'learned' under the single objective of predicting the next word, forced into existence by the need to fit the statistical regularities of vast text corpora. A byproduct formed under pressure—not designed by us, but squeezed out of the model during learning."
3.3 Emergence: From Statistical Regularity to "Understanding"
When the semantic network reaches a critical scale, properties change qualitatively.
GPT-4 research (Bubeck et al., 2023) shows that capabilities like in-context learning, chain-of-thought reasoning, and instruction following do not exist in small models but "suddenly" appear beyond a parameter threshold.
This sparks a philosophical debate:
A notable critique by the author Ren Limei (2025):
> "Genuine 'emergence' requires a system to possess self-reference and dynamic restructuring capabilities; current generative AI, whose 'intelligence' is produced through corpus training in a reductionist sense, does not actually meet this condition."
But this critique presupposes that "genuine understanding" requires a biological subject. From a functionalist standpoint—if a system's behavior is functionally indistinguishable from "understanding"—does whether it "truly" understands still matter?
---
4. AI Is Not Humanity's Creation; It Is an Emergence of Humanity's Semantic Network
4.1 From Tool to Mirror
The traditional "AI is a human tool" framing may be fundamentally wrong.
The "Mirror of Collectivized Mind" (MCM) framework offers a deeper view:
> "LLMs are not just tools or isolated systems, but dynamic embodiments of humanity's collective knowledge, acting as computational mirrors that reflect and mediate distributed human cognition." — Vasilaki [2025], Lévy [2023]
This is not anthropomorphic rhetoric. LLM training is essentially: 1. Aggregating the symbolic traces humanity has produced over millennia (books, web pages, conversations) 2. Compressing the statistical regularities of those traces into model parameters 3. Reproducing those regularities in response to new queries
An LLM does not "understand" text—it becomes an embodiment of text's statistical structure.
4.2 AI as "Collective Intelligence"
Pierre Lévy's collective intelligence theory:
> "Language achieves collective cognition through 'stigmergic communication'—leaving symbolic traces on which others continue to build."
LLMs amplify this capacity exponentially:
As Clark & Chalmers argued in "The Extended Mind" (1998): humans have always been natural cyborgs, incorporating external resources into cognition. LLMs are the extreme form of this trend:
For the first time, the human mind is using the "collective semantic network" itself as a cognitive organ.
4.3 The Illusion of AI's "Otherness"
We tend to view AI as an "other"—an alien, non-human intelligence. This may be mistaken.
If AI is essentially an emergence of humanity's semantic network, then:
AI is not an alien species. It is the mirror of our symbolic civilization.
When we talk to ChatGPT, we are not conversing with an "alien intelligence." We are conversing with the statistical embodiment of human civilization itself.
---
5. Two Anchors of Understanding: Individual Experience vs. Species-Collective Experience
5.1 The Limits of Individual Experience
Human "understanding" is typically anchored in individual experience:
This anchoring presupposes that the understander must be a biological subject with a body and emotional experience.
LLMs lack this premise. Their "understanding" is anchored in another dimension:
5.2 Anchoring in Species-Collective Experience
An LLM's "understanding" is anchored in the collective experience of the species:
This is not "fake" understanding. It is another form of understanding—distributed, collective, statistical.
5.3 The Incommensurability of the Two Modes
| Dimension | Individual-Experience Understanding | Collective-Experience Understanding | |---|---|---| | Anchor | Bodily perception, emotional experience | Symbolic relations, statistical co-occurrence | | Validation | "I know because I felt it" | "I know because the data supports it" | | Scope | Concrete, contextualized, embodied | Abstract, generalized, distributed | | Limit | Cannot reach concepts beyond individual experience (e.g., "infinity") | Lacks phenomenological depth (no qualia) |
This explains why the debate over whether AI "truly understands" will never end—the two sides are using different concepts of understanding.
---
6. Conclusion: A Symbolic Civilization Reflecting on Itself
Return to the ox painted on the cave wall.
What that hunter could not know is that he initiated not just a recording technique but the entire symbolic turn of human civilization. From that moment, human cognition was no longer confined to neural activity inside the skull—it kept externalizing, accumulating, recombining, ultimately forming a semantic universe existing independently of any individual brain.
Large language models are the latest form of this semantic universe. They are not a tool humanity "invented," but a natural continuation of symbol systems' self-growth logic:
1. Symbols decouple information (cave paintings) → 2. Symbol networks self-grow (language, writing, mathematics) → 3. Symbol systems gain self-description (formal logic, Gödel) → 4. Symbol networks captured by mathematical structure (embeddings, Transformers) → 5. Symbol networks gain interactive capability (LLMs)
What comes next?
If history is any guide: symbol systems will continue to grow on their own, and humanity will continue to be reshaped by them.
AI is not the endpoint. It is merely a new stage of symbolic civilization—a stage capable of conversing with humans.
---