🧭 Prelude: A Seven-Page Philosophical Bomb with No Code
In May 2026, Richard Sutton dropped a seven-page paper onto arXiv. Zero experiments, zero benchmark scores, not a single new algorithm. Yet this 2024 Turing Award winner, the widely recognized father of reinforcement learning, used pure philosophical argument to lay a foundation for the entire anti-large-model route.
After my first read, something felt off. The whole world was discussing "Sutton finally speaking out against large models," but nobody noticed what he actually wrote. Enactive Reinforcement Learning is not an evolution of generative AI—it stands squarely opposed to it. The paper's core is a single sentence: only by acting does the world reveal itself to you.
Yet this seemingly clear statement crashes directly into two iron pillars Sutton himself erected over thirty years ago. The first is called the Reward Hypothesis; the second, the Bitter Lesson. With both pillars breaking simultaneously, the foundation of a $5.1 billion wager begins to crack.
⚡ The True Meaning of Enactive: The World Is Not Painted, It Is Collided With
> A brief note on enactive cognition: This theory holds that cognition is not a matter of the brain first drawing a map of the world and then following it. It emphasizes that the body must act in real time, forming feedback loops with the environment, and the world gradually reveals itself within these loops. An infant reaching for a toy does not first model the laws of physics—it reaches, hits obstacles, adjusts, fails, and reaches again. The world only "manifests" through action. This is the opposite direction from large models' offline generation of static world representations through massive data.
Sutton invokes this concept to pass judgment on the current mainstream route. He believes the "generative" approach of large models is essentially still trying to stuff the world into parameters, violating the most primordial way intelligence arises. An agent must act first; only then does the world respond. A generative model is like a person sitting in a room forever drawing maps, while an enactive agent carries a lantern—each step illuminates a small patch of real ground.
This view sounds radical, and it effectively sentences today's hottest generative large models to death. The problem is that Sutton himself planted two time bombs under this verdict.
💰 A $5.1 Billion Wager and Three Tables of Chips
Sequoia, NVIDIA, and Google now sit at the same table. They have poured $1.1 billion into a company with zero products and zero revenue, driving its valuation straight to $5.1 billion—all betting that "Sutton is right."
The stakes are heavy for a simple reason: if large models truly are a dead end, then the next generation of embodied intelligence must take the enactive route—emphasizing real-time action, environment coupling, and online learning rather than offline pre-trained massive static models. The investors are betting on a fundamental paradigm shift in robotics and agents between 2028 and 2030.
But the foundation they are betting on is itself cracking.
🏛️ First Pillar: The Illusion of Autonomy and the Collapse of the Reward Hypothesis
Sutton's paper states in black and white: even within the enactive framework, normativity is still defined by external reward functions.
This sentence is equivalent to tearing down his own temple with his own hands. The Reward Hypothesis is one of Sutton's most central contributions—any intelligent goal can be reformulated as maximizing some scalar reward signal. What is most compelling about enactive cognition is precisely "autonomous emergence," yet the paper itself admits that what counts as "good" or "what should be done" is ultimately externally defined.
Imagine a dancer who claims to be free, only to discover that all the lights and music were pre-set by judges in the audience. Every seemingly improvised spin still chases that external carrot at its core. This is not autonomy—it is dancing in shackles. Sutton wants to raise the banner of enactive autonomy, only to find the flagpole deeply planted in the foundation of the Reward Hypothesis he laid himself. A temple of his own making, demolished by his own hand.
📜 Second Pillar: The Deadly Boomerang of the Bitter Lesson
Even more damning is the second blow.
Sutton has preached the Bitter Lesson for thirty years: hard-coding human prior knowledge into architectures is inferior to letting general methods plus massive computation grow on their own. Deep learning defeated all hand-crafted feature engineering precisely because it embraced scaling rather than hard-wiring specific theories.
Yet this time, Sutton hard-codes a "cognitive theory"—embodiment, real-time action-perception coupling, the enactive loop—directly into agent architecture. This is exactly the approach he once despised most: stuffing a particular human cognitive theory into the system instead of letting agents grow understanding from real interaction.
It is like an old captain who spent his life warning younger sailors "don't trust the compass; trust the real-time interplay of wind and sails," only to have the keel of his new ship engraved with "must strictly follow this compass." The lesson he preached for a lifetime has come back to slap him in the face.
👻 The Ghost of Brooks and the Same Play from Thirty Years Ago
Thirty years ago, Rodney Brooks proposed the subsumption architecture, arguing that intelligence needs no internal world representation and can emerge through layers of reactive behavior directly coupled with the environment. This closely matches the spirit of enactive cognition.
What happened? Deep learning came from behind and thoroughly defeated this "representation-free" route. Pure reactivity quickly hit a ceiling on complex long-horizon tasks, and scaling plus end-to-end learning swept everything aside. Brooks' philosophy lost to engineering reality.
Now Sutton has picked up a similar torch, but the engineering world has already voted with its actions. At ICLR 2026, Vision-Language-Action model submissions exploded from 9 the previous year to 164. Robots are already neatly folding clothes in laboratories. Nobody is waiting for the philosophy debate to finish—they only care about which method runs faster and more stably on real physical tasks.
🧩 A Thirty-Year-Old Unsettled Account from Cognitive Science
Sutton's paper has inadvertently stepped on two old blades of cognitive science.
One is the scaling-up problem: how do low-level action-perception loops give rise to abstract reasoning, language planning, and long-term memory? Enactive approaches may work in simple scenarios, but on tasks requiring complex cognition, will they repeat the fate of Brooks' route?
The other is the coupling-constitution fallacy: does interaction between environment and body merely "couple with" cognition, or does it genuinely "constitute" part of cognition? This debate has raged for thirty years without resolution. The paper dredges up this old account but offers no new answer.
> Extended note: These debates are not academic parlor games. They directly determine how much world knowledge should be packed into models today, and how much should be left to real-time interactive learning. If the constitution thesis prevails, we may need entirely new hardware-software co-designed architectures; if it is mere coupling, the current VLA route may still have years of fight left in it. Sutton has pushed this undecided battlefield back in front of everyone.
⏳ The Countdown of Three Tables of Bets
The story ends with three different tables of chips.
The first table bets on 2028: believing scaling will soon hit a wall, and the pure enactive route will dominate next-generation embodied intelligence.
The second table also looks to 2028, but bets on a hybrid route—large models provide the prior skeleton, while enactive methods handle online adaptation and error correction.
The third table bets on 2030: believing the engineering faction will keep pushing the current paradigm to its extreme with scaling and data, and the philosophical debate will be a mere historical footnote.
Which table's chips will remain on the felt? The answer will not be argued out in arXiv comment sections. It will only be decided in the real commercial battlefield of the late 2020s—with real money, real robots, and real user experience.
🏁 Conclusion: Philosophy Is the Map; the Battlefield Is the Touchstone
Sutton's seven-page paper is both a prescient warning and a blueprint riddled with self-contradiction. It reminds us that intelligence may fundamentally reside in the dance between action and environment, not in world models piled into parameters. But it also exposes that even the wisest minds can collide with walls they themselves built inside the philosophical labyrinth.
Having studied reinforcement learning for twenty years, I have witnessed too many paradigm shifts. Every time, someone declares "this time is different," and every time, engineering reality drags philosophy back to earth with cold, hard metrics. Sutton's self-contradiction this time is worth $5.1 billion. Who ultimately foots that bill depends on how the real world responds.
Philosophy matters. It helps us see the direction. But what ultimately decides AI's future is never logical self-consistency in papers—it is real performance running in products. The shrine has developed cracks. Whether it gets repaired or collapses entirely, the commercial battlefield of 2028 to 2030 will provide the answer.
📚 References
1. Sutton, R. S. (2026). Enactive Reinforcement Learning: A Philosophical Position Paper. arXiv preprint. 2. Sutton, R. S. (2019). The Bitter Lesson. Incomplete Ideas. 3. Brooks, R. A. (1991). Intelligence without representation. Artificial Intelligence, 47(1-3), 139-159. 4. Varela, F. J., Thompson, E., & Rosch, E. (1991). The Embodied Mind: Cognitive Science and Human Experience. MIT Press. 5. Industry reports and ICLR 2026 submission trends on Vision-Language-Action models.