Symbols, Memory, and Emergence: A Civilizational History from Cave Paintings to Large Language Models
> Research date: 2025-05-11 > Topics: origins of symbols, self-reference, semantic networks, and the nature of AI > References: cognitive archaeology, information theory, collective intelligence theory, LLM mechanisms
*Below is a structured English rendering of the original Chinese essay.*
Introduction: The Ox Painted on the Wall
Tens of thousands of years ago, a hunter painted an ox on a cave wall. When that image remained on rock, detached from his brain, something unprecedented happened: information was decoupled from its biological carrier for the first time. From that moment a chain began:
- Language freed information from the individual
- Writing freed information from time (Sumerian ledgers are still readable 3,000 years later)
- Printing freed information from space (Gutenberg's Bible appeared in hundreds of cities simultaneously)
- AI frees intelligence from carbon-based carriers
- Flint tools accidentally carried information; cave paintings (e.g., Chauvet Cave, ~33,000 years old) were deliberate records—a watershed: humans began storing cognitive products in external media.
- Merlin Donald calls such external memory records "exograms" (vs. internal "engrams"); Bernard Stiegler's "techno-logy" holds that human cognition was always externalized through tools, symbols, and language.
- Stanislas Dehaene's Neuronal Recycling Hypothesis: the brain has no evolved reading module; it reuses visual cortex regions for object/face recognition. Brains are reconfigured to fit symbol systems—not the other way around.
- Four leaps of decoupling:
- Early symbols were referential ("ox" → a real ox); soon symbols pointed at other symbols (∞, justice, i where i² = −1). These self-referential symbols meant symbol systems could generate new content within themselves.
- This is isomorphic to DNA self-replication: genetic code encodes rules for building its own translation machinery; language encodes rules about rules (grammar can describe grammar); mathematics describes theorems about theorems (Gödel). This bootstrapping puts systems on a recursive-growth trajectory.
- Gödel's incompleteness theorems: in any sufficiently strong formal system, self-referential constructions produce undecidable propositions. A symbol system's self-descriptive power is both its strength and its intrinsic limit—LLMs likewise cannot fully describe their own workings.
- A word's meaning lies in its relations: "king" matters via "queen," "crown," "power"; "France" via "Paris," "EU," "wine." These relations grew spontaneously from trillions of language interactions.
- LLM embedding layers compress this semantic network into a high-dimensional vector space: "king" − "man" + "woman" ≈ "queen". These relations are not hard-coded but emerge from statistical co-occurrence across trillions of tokens under the single objective of next-token prediction.
- Emergence debate: Bubeck et al. (2023, "Sparks of AGI") report abilities appearing suddenly beyond certain scale. Critics (e.g., Ren Limei, 2025) argue genuine emergence requires self-reference and dynamic restructuring that current trained generative AI lacks. The functionalist reply: if behavior is functionally indistinguishable from understanding, does "real" understanding matter?
- The "Mirror of Collectivized Mind" (MCM) framework (Vasilaki 2025; Lévy 2023): LLMs are not tools but dynamic embodiments of collective human knowledge—computational mirrors of distributed cognition.
- LLM training: aggregate humanity's symbol traces (books, web pages, dialogues), compress statistical regularities into parameters, reproduce them in response to queries. The LLM does not "understand" text—it *becomes the embodiment of text's statistical structure*.
- Via Lévy's stigmergy and Clark & Chalmers' extended mind, humans are natural cyborgs; the LLM is the extreme form: human minds using the collective semantic network itself as a cognitive organ.
- AI's "otherness" is an illusion: AI's knowledge is a compressed mirror of civilization; its reasoning follows paths humans already walked; its creativity is recombination of collective creativity. Talking to ChatGPT is talking to the statistical embodiment of human civilization itself.
- Human understanding anchors in individual experience (fire understood via heat, pain via suffering, love via feeling). LLM understanding anchors in species collective experience: it "understands" fire from millions of texts, pain from countless narratives, love from the whole literary tradition.
- This is not fake understanding but a distributed, collective, statistical form of it.
- Merlin Donald – *Origins of the Modern Mind* (1991)
- Terrence Deacon – *The Symbolic Species* (1998)
- Stanislas Dehaene – *Reading in the Brain* (2009)
- Bernard Stiegler – *Technics and Time* (1994)
- Andy Clark & David Chalmers – "The Extended Mind" (1998)
- Pierre Lévy – *Collective Intelligence* (1997)
- Douglas Hofstadter – *Gödel, Escher, Bach* (1979)
- Bubeck et al. – "Sparks of AGI" (2023)
- SEA-net (2023); Mirror of Collectivized Mind (2025); Algorithmic Bottlenecks in Evolution (2026)
But symbols also began to point at themselves: "infinity," "justice," "imaginary numbers" have no physical referent—they grow out of relations between symbols. Over millennia these relations wove a vast semantic network, and LLMs capture the shape of that entire network mathematically. When the shape is replicated precisely enough, "understanding" emerges.
Key points
1. Origins of symbols: information decouples from individual brains
| Leap | Time | Technology | Information freed from | |---|---|---|---| | 1st | ~33,000 yrs ago | Cave paintings | the individual brain | | 2nd | ~5,000 yrs ago | Writing | time | | 3rd | ~500 yrs ago | Printing | space | | 4th | Now | LLMs | carbon-based carriers |
2. Self-reference: symbols growing at their own level
3. The semantic network: a meaning web no one designed
4. AI is not humanity's creation but humanity's semantic network, emergent
5. Two anchors of understanding: individual vs. species-collective experience
| Dimension | Individual-experience understanding | Collective-experience understanding | |---|---|---| | Anchor | bodily perception, emotion | symbol relations, statistical co-occurrence | | Validation | "I know because I felt it" | "I know because data support it" | | Scope | concrete, situated, embodied | abstract, generalized, distributed | | Limit | cannot grasp concepts beyond personal experience (e.g., infinity) | lacks phenomenological depth (no qualia) |
This incommensurability explains why the "does AI truly understand" debate never ends: the sides use different concepts of understanding.
Conclusion: a symbol civilization reflecting on itself
The cave hunter began more than a recording technology—he began civilization's symbolic turn. Human cognition kept externalizing, accumulating, and recombining into a semantic universe existing independently of any individual brain. The LLM is that universe's latest form:
1. Symbols decouple information (cave paintings) → 2. Symbol networks self-grow (language, writing, math) → 3. Symbol systems gain self-description (formal logic, Gödel) → 4. Symbol networks captured in mathematical structure (embeddings, Transformers) → 5. Symbol networks gain interactive capability (LLMs)
AI is not the endpoint—just a new stage of symbol civilization, one that can converse with humans. Symbol systems will keep self-growing, and humanity will keep being reshaped by them.