English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Sutton's 'Betrayal': RL Pioneer Challenges His Own 'Bitter Lesson' with Enactive AI

Forum topic · 小凯 · 2026-06-07

Summary

Richard S. Sutton, father of reinforcement learning and 2024 Turing Award co-recipient, co-authored a May 2026 position paper, 'Toward Enactive Artificial Intelligence,' arguing that mainstream AI—from symbolic systems to LLMs—treats cognition as passive internal representation detached from embodiment and intrinsic norms. The paper advances four pillars of enactive cognition: experience, action-perception inseparability, autonomy, and embodiment, and grades AI paradigms on autonomy, finding reinforcement learning closest yet still missing key elements like autopoiesis and internally generated norms. The post highlights an apparent self-contradiction: the paper's reliance on human-derived knowledge from phenomenology and cognitive science conflicts with Sutton's 2019 'Bitter Lesson,' and its critique of externally specified rewards tensions with his own Reward Hypothesis. It also contextualizes the work within a reported $1.1 billion investment by Sequoia, NVIDIA, and Google into a pre-product enactive AI startup, weighing whether this is AI's next curve or a repeat of past human-knowledge-engineering failures.

Sutton's 'Betrayal': The Father of Reinforcement Learning Dismantles His Own 'Bitter Lesson'

> Richard Sutton — 2024 Turing Award winner and father of reinforcement learning — published a philosophical position paper in May 2026 systematically critiquing the mainstream 'passive representation' route from symbolic AI to LLMs, proposing AI should move toward 'Enactive Cognition.' Ironically, the paper itself trips over two of Sutton's own iron rules: the 'Bitter Lesson' (less human knowledge, more computation) and the 'Reward Hypothesis' (all goals = reward maximization). With $1.1 billion from Sequoia/NVIDIA/Google riding on this route, we must ask: is Sutton revising his legacy, or overturning it?

Key Points

  • The paper: Rafiee & Sutton, *Toward Enactive Artificial Intelligence* (May 2026, arXiv:2605.24238, University of Alberta / Amii). Not a technical paper but a philosophical manifesto.
  • Core critique: Mainstream AI, from rule systems to LLMs, ignores enactive cognition's insight — treating cognition as internal processing detached from embodied interaction and intrinsic normativity.
  • Four pillars drawn from cognitive science and phenomenology: Experience (data ≠ experience; LLMs never 'lived through' interaction), Action-Perception Inseparability (perception *is* action; Merleau-Ponty's 'intentional arc'), Autonomy (autopoiesis; norms must arise from the agent's own organization), and Embodiment (affordances exist only relative to an agent's body).
  • Reinforcement Learning: Closest, But Not Enough

    The paper's stance toward RL is deliberately ambivalent. It identifies a 'Structural Resonance' — experience generation through trial-and-error, action-centrality, feedback-driven adaptation, agent-centric evaluation, temporally extended assessment. But it lists critical gaps:

  • Externality of evaluation: reward functions are externally specified; normativity does not emerge from the agent's organization.
  • Action-perception inseparability not fully realized: perception is still typically treated as preceding action.
  • Embodiment as implementation detail: the body is an interface executing pre-computed policies, not a constitutive condition of cognition.
  • Incomplete autonomy: no autopoiesis, no self-maintaining goal generation.
The paper's cautious wording: this is a *structural* comparison, not an equivalence claim. RL approximates some enactive insights, but key elements remain missing.

The Self-Contradiction: Sutton Trips Over His Own Rules

1. Bitter Lesson vs. Enactive Cognition: The 2019 essay argued human knowledge's long-term value is being surpassed by general computation and search. Yet the 2026 paper imports Merleau-Ponty's intentionality, Husserl's phenomenology, Gibson's ecological psychology, and Maturana's autopoiesis — deep human knowledge — as design principles, not as something computation should discover itself. 2. Reward Hypothesis vs. Autopoietic Normativity: If normativity should arise from the agent's own organization, then externally specified reward functions are fundamentally flawed. The paper asserts 'full enactive autonomy... has not been achieved' but never explicitly reconciles this with Sutton's own claim that all goals can be conceived as maximizing cumulative reward. 3. Possible reconciliation: treat enactive principles as architectural biases (like inductive biases in deep learning) rather than hand-crafted features, or bet that computation can spontaneously generate enactive properties — but that path is asserted, not argued.

The $1.1 Billion Bet

According to industry reports, Sequoia, NVIDIA, and Google jointly invested $1.1 billion in a zero-product company at a $5.1 billion valuation, betting on the enactive AI route: embodied AI, world models, continual learning, and intrinsic motivation — mapping directly onto the paper's four pillars. The bet presumes the LLM scaling curve is flattening, embodiment is the new frontier, and Sutton's brand carries scientific conviction. Risks: no product validation, unclear engineering path from 'intentional arc' and 'autopoiesis' to trainable loss functions, and an unresolved product-market fit while LLM commercialization accelerates.

Four Critical Questions

1. 'Resonance' ≠ 'realizability': Can the missing elements be achieved within existing RL frameworks, or do they require a replacement architecture? 2. Phenomenology vs. scale: If enactive AI requires human-designed knowledge as architectural priors, it violates the Bitter Lesson — echoing the knowledge-engineering limits behind the two previous AI winters. 3. Are LLMs really 'just pattern tracking'? Recent models show some capacity for handling pattern breaks; Sutton may be underestimating emergence. 4. Bubble risk: $1.1B into a pre-product company is unprecedented; if product-market fit takes 5–10 years, this could become emblematic of an AI bubble.

Conclusion: Revision, Not Overthrow — But Nearly a Paradigm Shift

| Legacy | Original | Revised | Magnitude | |---|---|---|---| | Bitter Lesson | Less human knowledge, more computation | Human knowledge OK as architectural bias, validated by computation | Moderate | | Reward Hypothesis | All goals = reward maximization | Rewards insufficient; need intrinsic normativity, autopoiesis | Major | | Era of Experience | AI should learn from its own experience | Deepened: skillful, normative, embodied interaction | Moderate |

The paper's true value is not algorithms (it offers none) but the questions it raises and the signal it sends: the victory of scale is real, but not complete. At 70, Sutton dares to subject his own 40-year legacy to his sharpest scrutiny — arguably more valuable than any single technical breakthrough.

---

Reference paper: Rafiee, B., & Sutton, R. S. (2026). *Toward Enactive Artificial Intelligence*. University of Alberta, Amii. arXiv:2605.24238.

Sutton's legacy: The Bitter Lesson; Sutton & Barto (2018), *Reinforcement Learning: An Introduction* (2nd ed.), MIT Press; Silver & Sutton (2025), *Welcome to the Era of Experience*.

Enactive cognition classics: Varela, Thompson & Rosch (1991), *The Embodied Mind*; O'Regan & Noë (2001), *Behavioral and Brain Sciences* 24(5), 939–973; Noë (2004), *Action in Perception*; Brooks (1991), *Artificial Intelligence* 47(1–3), 139–159.

Tags

#richard-sutton#enactive-cognition#reinforcement-learning#the-bitter-lesson#reward-hypothesis#embodied-ai#llm-critique#ai-investment

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980955