English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

UC San Diego Scholars Argue in Nature That LLMs Like GPT-4.5 Already Constitute AGI

Forum topic · 小凯 · 2026-06-11

Summary

Four UC San Diego scholars—Eddy Keming Chen (philosophy), Mikhail Belkin (machine learning), Leon Bergen (linguistics), and David Danks (data science)—published a Nature comment arguing that, by reasonable standards, current large language models already constitute artificial general intelligence (AGI). Their central evidence: GPT-4.5 was judged human 73% of the time in a 2025 Turing test, fulfilling Alan Turing's 75-year-old 'imitation game' vision, alongside broad cross-domain competence in math, coding, science, and writing, and novel problem-solving with cross-domain transfer. The authors analyze three psychological mechanisms behind resistance to this conclusion: unreasonably high standards (perfection, omniscience, human-likeness, superintelligence), emotional resistance and conceptual confusion between intelligence and consciousness, and practical anxieties about jobs, education, and governance. They address counterarguments from the ARC-AGI benchmark, where frontier models score below 50% versus human 85%+, and brittleness critiques, arguing critics conflate imperfection with non-generality. The paper's significance lies in shifting the agenda from building AGI to managing systems that already exist—governance, education redesign, equitable benefit distribution, and preventing misuse.

UC San Diego Scholars Argue in Nature That LLMs Like GPT-4.5 Already Constitute AGI

In one sentence: Four UC San Diego scholars from philosophy, machine learning, linguistics, and cognitive science published a Nature comment arguing that large language models (like GPT-4.5), defined by reasonable standards of "general intelligence," already constitute AGI. The core evidence is not a single benchmark but the realization of Alan Turing's 75-year-old "imitation game" vision—GPT-4.5 was judged human by 73% of interrogators in a Turing test. The article's real value lies not in declaring victory, but in dissecting three psychological mechanisms behind why people refuse to accept that AGI has arrived.

Paper Details

| Dimension | Content | |---|---| | Title | Does AI already have human-level intelligence? The evidence is clear | | Authors | Eddy Keming Chen (philosophy), Mikhail Belkin (ML/CS), Leon Bergen (linguistics/CS), David Danks (data science/philosophy/policy) | | Institution | UC San Diego | | Published | Nature 650:36-40 (2026-02-02) | | Citations | 18 (as of time of search) | | Core conclusion | By reasonable standards, current LLMs already constitute AGI |

Core Argument: Three Lines of Evidence

Evidence 1: The Turing Test Has Been Passed

Another UC San Diego research group found in March 2025 that GPT-4.5 was judged to be human in 73% of cases in a Turing test—far exceeding the rate at which real humans are correctly identified. Alan Turing's 1950 "imitation game" asked: if a machine can converse via text in a way indistinguishable from a human, should we grant it intelligence? Seventy-five years later, that threshold has been crossed.

Evidence 2: Cross-Domain General Capability

GPT-4.5 shows high generality across:

  • Mathematics: complex reasoning and proof
  • Programming: code generation, debugging, optimization
  • Scientific reasoning: interdisciplinary problem-solving
  • Writing: creative, academic, and technical writing
  • Multilingual: cross-lingual transfer and translation
  • The key point is not that every domain reaches top human level, but that a single system achieves competent performance across such diverse domains—itself evidence of generality.

    Evidence 3: Novel Problem-Solving and Transfer

    The authors stress these systems are not merely "repeating training data." They can:

  • Solve problems not explicitly seen during training
  • Transfer knowledge from one domain to another
  • Perform compositional reasoning, producing outputs never seen in that form in training data
  • Why Do People Refuse to Accept It? Three Psychological Mechanisms

    Mechanism 1: Standards Set Unreasonably High

    Demands for "general intelligence" often include:

  • Perfection: never erring (but humans err too)
  • Omniscience: knowing everything (but humans don't)
  • Human-likeness: requiring emotion, consciousness, a body (but intelligence is functional, not tied to a substrate)
  • Superintelligence: exceeding humans (but AGI means "human-level," not "superhuman")
  • The authors note these added requirements are unreasonable—by the same standards, many humans would fail to qualify as having general intelligence.

    Mechanism 2: Emotional Resistance and Conceptual Confusion

  • Emotional resistance: accepting that AGI has arrived means rethinking human uniqueness, the value of work, and social structures—producing existential anxiety
  • Conceptual confusion: conflating "intelligence" with "consciousness," "self-awareness," and "soul." The article explicitly distinguishes: intelligence is a functional concept (the ability to solve problems), not an ontological one (the essence of being)
  • Mechanism 3: Practical Anxiety

    If AGI already exists, then:

  • What happens to labor markets?
  • Does education need complete restructuring?
  • Are existing AI governance frameworks adequate?
  • Where are humanity's anchors of value?
  • These concerns push people toward delaying acknowledgment of AGI's arrival to buy preparation time.

    Key Conceptual Clarifications: Four "Does Not Equal"

    | Common Misconception | Correct Understanding | |---|---| | General intelligence = perfection/omnicompetence | General intelligence = competent performance across broad domains, with room for error | | General intelligence = human-like | Intelligence is a function, not tied to biological bodies or self-awareness | | General intelligence = superintelligence | AGI = human-level; ASI = superhuman-level—two different stages | | Immediate economic disruption = the AGI standard | A lag exists between a technology's existence and its socioeconomic impact, historically normal |

    Counterarguments: ARC-AGI and the Brittleness Critique

    The article does not ignore dissent. The most systematic objections come from:

    The ARC-AGI benchmark (François Chollet):

  • Measures "fluid intelligence"—solving genuinely novel problems with minimal data
  • Current frontier models score <50%, humans 85%+
  • The $1 million prize remains unclaimed
  • Brittleness:

  • Systems excel at complex scientific reasoning
  • Yet make elementary errors on unexpected variants of simple tasks (e.g., counting letters)
  • This shows a gap between "surface capability" and "deep understanding"
  • How does the article respond? The authors argue these critiques conflate "imperfect" with "not general." Humans also fail on ARC-AGI and also have brittle moments. The definition of general intelligence should not require perfection.

    The Paper's Real Significance: Redefining the Agenda, Not Declaring Victory

    If one accepts the "AGI is here" premise, the agenda fundamentally shifts:

    | Old Agenda (AGI is future) | New Agenda (AGI is present) | |---|---| | Pursue breakthroughs to reach AGI | Manage risks and impacts of existing AGI systems | | Research "how to build AGI" | Research "how to coexist with AGI" | | Ethics is forward-looking | Ethics is urgent, real-time policy | | Governance frameworks can wait | Governance must catch up immediately |

    The conclusion is clear: whether or not you accept the "AGI" label, the capability level of current systems forces us to think about risk, governance, and coexistence with entirely new frameworks.

    Controversy and Reception

    The comment has sparked ongoing academic debate:

  • Supporters: say it clarifies conceptual confusion, moving discussion from "semantic disputes" to "capability assessment"
  • Opponents: say it lowers the AGI bar and "weaponizes definition"—letting tech companies claim AGI earlier for capital and policy advantages
  • Middle ground: acknowledging unprecedented LLM generality while insisting the "human-level" benchmark needs stricter empirical validation
  • An interesting data point: this Nature comment garnered 18 citations within 4 months, showing it has genuinely triggered scholarly discussion.

    Conclusion: The Label Matters Less Than Action

    The most valuable part of the article is its pragmatic pivot: whether or not you use the word "AGI," the capability reality of current systems has changed everything. Rather than debating labels, the priorities are:

    1. Build governance frameworks matching system capabilities 2. Redesign education and employment systems 3. Ensure equitable distribution of the technology's benefits 4. Prevent misuse or concentration of capabilities

    As the authors put it: "Eyes unclouded by anxiety can see the evidence clearly."

    References

  • Chen, E. K., Belkin, M., Bergen, L., & Danks, D. (2026). *Does AI already have human-level intelligence? The evidence is clear*. Nature, 650, 36-40. https://www.nature.com/articles/d41586-026-00285-6
  • UC San Diego Today coverage: https://today.ucsd.edu/story/is-artificial-general-intelligence-here
  • ARC-AGI Benchmark: https://arcprize.org/arc-agi
  • ResearchGate discussion: https://www.researchgate.net/publication/400368037

Tags

#agi#nature#gpt-4.5#turing-test#large-language-models#philosophy-of-mind#ai-governance#cognitive-science

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981100