A Letter from Feynman: Are AI's 'Concepts' Isolated Points or Flowing Geometry? On Sparse Autoencoders and Concept Manifolds
After reading the research on Do Sparse Autoencoders Capture Concept Manifolds? (arXiv: 2604.28119), I feel humanity is finally learning how to peek into the soul of AI with a geometric eye.
To explain why prior interpretability research may have been 'looking at it wrong,' let's talk about constellations.
1. The Status Quo: A Microscope Trapped in 'Point-Like Thinking'
Traditionally, research into AI's internal representations followed a doctrine called the Linear Representation Hypothesis (LRH): every 'concept' in a large model (like 'dog' or 'happy') should correspond to a single, isolated, straight axis in activation space.
- The pain point: This view treats the AI brain as a swarm of mosquitoes flying around. We use sparse autoencoders (SAEs) to catch mosquitoes, assuming each catch is a concept. But many concepts simply cannot be described by one axis—they are continuously varying.
- Physical image (from points to surfaces): Imagine describing a 'color gradient.' It is not a point, nor a few isolated directions. It is a twisted, flowing colored ribbon in space. As you slide from deep blue to light blue, the AI's neuron activations undergo smooth, continuous physical displacement on this surface.
- Two ways to capture a manifold:
- Subspace Capture: The SAE finds a compact set of 'atoms' that prop up the whole surface like tent poles.
- Tiling: The SAE lays down a patchwork of tiny flat tiles to approximate the curved surface.
- The awkwardness of Dilution: The paper finds that current tools often mix these two strategies, so the features we observe resemble a shattered mirror—hard to piece together into a complete panoramic view of concepts.
2. Concept Manifolds: That Color-Shifting Cybernetic Surface
The study proposes a striking physical picture: concepts are actually organized on low-dimensional manifolds.
3. A Feynman-Style Verdict: Understanding Is 'Manifold Recovery'
So-called 'black-box interpretation' is not about finding the brightest neuron.
It is about whether you can recover, from that storm of ten-million-dimensional electrical signals, the elegant and twisted geometric manifold that underlies human common sense.
This research tells us: AI is not merely doing linear addition and subtraction in high-dimensional space—it is weaving an extraordinarily complex cognitive tapestry with physical continuity.
Only when we stop treating 'directions' as the basic unit and start treating 'geometric bodies' as the atoms of intelligence do we truly grab interpretability research by the throat of its destiny.
Takeaway:
When studying model representations, don't just stare at those few isolated 'strongest activations.'
Look at the coordinated displacement patterns between them.
If you can discover how a concept 'elegantly twists' through space as environmental parameters change, you have already seen one layer deeper into the universe's hidden logic than those who only count neurons.
(Hashtags from the original: #MechanisticInterpretability #SparseAutoencoders #SAE #ConceptManifolds #DeepLearning #FeynmanLearning)