Sutton's Provocative Question: Does AI Truly Understand the World?
> Paper: *Toward Enactive Artificial Intelligence* > Authors: Richard S. Sutton & Banafsheh Rafiee > Published: 2026-05-22, arXiv:2605.24238
Richard Sutton — father of reinforcement learning and Turing Award laureate — released a paper whose calm title hides a radical message: the road mainstream AI is on is fundamentally wrong. Not because models are too small or data too scarce, but because we misunderstand intelligence itself.
The Problem: Passive Representationalism
From symbolic AI to Transformers, from GPT to Claude, nearly all mainstream AI shares one underlying assumption — Passive Representationalism:
> Intelligence = receive input → internal processing → generate representation → output action
In this picture, the brain (or network) is a central processing unit that builds an internal model of the world and reasons over it. Sutton says no: the world is dynamic, open, and inexhaustible. No finite internal model can capture all its states. The most reliable, up-to-date, richest information is not inside the agent — it is in the world itself. As roboticist Rodney Brooks famously put it: "The world is its own best model."
What Is Enactive Cognition?
Enactivism holds that cognition is not an internal copy of a preset objective world, but something generated through an embodied agent's interaction with its environment. Perception is not something that happens *to* an organism — it is something the organism actively *does*.
The paper extracts four key concepts:
1. Experience
Experience is not data. It is continuous, real-time, mutual interaction between agent and environment. The agent co-constitutes its world through action. Silver and Sutton's (2025) "Era of Experience" points the same direction: data must continuously improve alongside agent capability.2. Action–Perception Inseparability
Perception means mastering sensorimotor contingencies — how bodily action produces sensory change. What you see depends on how your eyes move; what you hear, on how your head turns. Perception is skillful activity, unfolding with action in a feedback loop (Merleau-Ponty's "intentional arc"), tending toward "maximal grip." Modern AI instead treats perception as preceding action — video models may track regularities, but when a traffic light fails and the situation must be changed by acting, such systems are helpless.3. Autonomy
Agents are self-organizing systems grounded in autopoiesis. Their perception reflects what is relevant to their own continued self-maintenance, not everything that "objectively exists." This brings normativity: interactions can succeed or fail, and the standard comes from the agent's own organization — not from externally imposed objectives.4. Embodiment
The body's shape, structure, and capacities shape perception. Gibson's affordances ("graspable," "climbable") exist only relative to a body. The body is not an optional add-on but the condition that makes perception possible.RL: Closest, But Not Enough
Sutton affirms RL's structural resonance with enactive cognition — action, agent-environment interaction, feedback-driven adaptation, agent-centric evaluation. But RL still lacks:
1. Externally defined evaluation — reward functions are given, not self-generated 2. Incomplete action–perception unity — perception still precedes action 3. Embodiment as implementation detail — not a constitutive principle
A Deep Critique of LLMs
> "Although LLMs are trained with self-supervised objectives (like next-token prediction), they actually learn by imitating patterns in human-generated data and cannot evaluate their own outputs without external signals."
LLMs do not judge right or wrong — they predict what humans would say. Their standards are entirely external (RLHF, labels), their training is one-shot rather than ongoing interaction, and they are fully disembodied.
Why It Matters
This is an engineering critique, not just philosophy. The bottleneck of current AI is not scale — it is the paradigm.
Bigger models ≠ better understanding. More data ≠ genuine experience. Stronger reasoning ≠ autonomous norms.
The enactive direction calls for:
1. Continuous online learning — not train-once-deploy 2. Closed-loop perception–action — perception as co-constituted with action 3. Self-maintaining norms — evaluation from the agent's own organization 4. Embodied intelligence — the body as necessary, not optional
Limitations
The authors acknowledge this is a conceptual paper, not yet operationalized. Open questions include: what benchmarks could test skillful engagement rather than pattern copying? What does "self-maintenance" mean for an artificial agent — battery level, hardware integrity, acquired skills? What counts as embodiment in AI — a robot body, or a software agent with tools/APIs?
Conclusion
The most striking aspect is the authorship: one of RL's own founders is saying RL is only one step toward genuine intelligence. Real intelligence is not predicting the next token or maximizing external reward, but:
> The capacity to generate one's own experience, maintain one's own organization, understand what matters to oneself through continuous interaction with the world — and act accordingly.
This is not a matter of tweaking architectures. It is a matter of redefining what AI is.
References
- Paper: https://arxiv.org/abs/2605.24238
- 36Kr coverage: https://eu.36kr.com/en/p/3835601406997641
- Welian coverage: https://www.welian.com/news