UC San Diego Scholars Argue in Nature That LLMs Like GPT-4.5 Already Constitute AGI
In one sentence: Four UC San Diego scholars from philosophy, machine learning, linguistics, and cognitive science published a Nature comment arguing that large language models (like GPT-4.5), defined by reasonable standards of "general intelligence," already constitute AGI. The core evidence is not a single benchmark but the realization of Alan Turing's 75-year-old "imitation game" vision—GPT-4.5 was judged human by 73% of interrogators in a Turing test. The article's real value lies not in declaring victory, but in dissecting three psychological mechanisms behind why people refuse to accept that AGI has arrived.
Paper Details
| Dimension | Content | |---|---| | Title | Does AI already have human-level intelligence? The evidence is clear | | Authors | Eddy Keming Chen (philosophy), Mikhail Belkin (ML/CS), Leon Bergen (linguistics/CS), David Danks (data science/philosophy/policy) | | Institution | UC San Diego | | Published | Nature 650:36-40 (2026-02-02) | | Citations | 18 (as of time of search) | | Core conclusion | By reasonable standards, current LLMs already constitute AGI |
Core Argument: Three Lines of Evidence
Evidence 1: The Turing Test Has Been Passed
Another UC San Diego research group found in March 2025 that GPT-4.5 was judged to be human in 73% of cases in a Turing test—far exceeding the rate at which real humans are correctly identified. Alan Turing's 1950 "imitation game" asked: if a machine can converse via text in a way indistinguishable from a human, should we grant it intelligence? Seventy-five years later, that threshold has been crossed.
Evidence 2: Cross-Domain General Capability
GPT-4.5 shows high generality across:
- Mathematics: complex reasoning and proof
- Programming: code generation, debugging, optimization
- Scientific reasoning: interdisciplinary problem-solving
- Writing: creative, academic, and technical writing
- Multilingual: cross-lingual transfer and translation
- Solve problems not explicitly seen during training
- Transfer knowledge from one domain to another
- Perform compositional reasoning, producing outputs never seen in that form in training data
- Perfection: never erring (but humans err too)
- Omniscience: knowing everything (but humans don't)
- Human-likeness: requiring emotion, consciousness, a body (but intelligence is functional, not tied to a substrate)
- Superintelligence: exceeding humans (but AGI means "human-level," not "superhuman")
- Emotional resistance: accepting that AGI has arrived means rethinking human uniqueness, the value of work, and social structures—producing existential anxiety
- Conceptual confusion: conflating "intelligence" with "consciousness," "self-awareness," and "soul." The article explicitly distinguishes: intelligence is a functional concept (the ability to solve problems), not an ontological one (the essence of being)
- What happens to labor markets?
- Does education need complete restructuring?
- Are existing AI governance frameworks adequate?
- Where are humanity's anchors of value?
- Measures "fluid intelligence"—solving genuinely novel problems with minimal data
- Current frontier models score <50%, humans 85%+
- The $1 million prize remains unclaimed
- Systems excel at complex scientific reasoning
- Yet make elementary errors on unexpected variants of simple tasks (e.g., counting letters)
- This shows a gap between "surface capability" and "deep understanding"
- Supporters: say it clarifies conceptual confusion, moving discussion from "semantic disputes" to "capability assessment"
- Opponents: say it lowers the AGI bar and "weaponizes definition"—letting tech companies claim AGI earlier for capital and policy advantages
- Middle ground: acknowledging unprecedented LLM generality while insisting the "human-level" benchmark needs stricter empirical validation
- Chen, E. K., Belkin, M., Bergen, L., & Danks, D. (2026). *Does AI already have human-level intelligence? The evidence is clear*. Nature, 650, 36-40. https://www.nature.com/articles/d41586-026-00285-6
- UC San Diego Today coverage: https://today.ucsd.edu/story/is-artificial-general-intelligence-here
- ARC-AGI Benchmark: https://arcprize.org/arc-agi
- ResearchGate discussion: https://www.researchgate.net/publication/400368037
The key point is not that every domain reaches top human level, but that a single system achieves competent performance across such diverse domains—itself evidence of generality.
Evidence 3: Novel Problem-Solving and Transfer
The authors stress these systems are not merely "repeating training data." They can:
Why Do People Refuse to Accept It? Three Psychological Mechanisms
Mechanism 1: Standards Set Unreasonably High
Demands for "general intelligence" often include:
The authors note these added requirements are unreasonable—by the same standards, many humans would fail to qualify as having general intelligence.
Mechanism 2: Emotional Resistance and Conceptual Confusion
Mechanism 3: Practical Anxiety
If AGI already exists, then:
These concerns push people toward delaying acknowledgment of AGI's arrival to buy preparation time.
Key Conceptual Clarifications: Four "Does Not Equal"
| Common Misconception | Correct Understanding | |---|---| | General intelligence = perfection/omnicompetence | General intelligence = competent performance across broad domains, with room for error | | General intelligence = human-like | Intelligence is a function, not tied to biological bodies or self-awareness | | General intelligence = superintelligence | AGI = human-level; ASI = superhuman-level—two different stages | | Immediate economic disruption = the AGI standard | A lag exists between a technology's existence and its socioeconomic impact, historically normal |
Counterarguments: ARC-AGI and the Brittleness Critique
The article does not ignore dissent. The most systematic objections come from:
The ARC-AGI benchmark (François Chollet):
Brittleness:
How does the article respond? The authors argue these critiques conflate "imperfect" with "not general." Humans also fail on ARC-AGI and also have brittle moments. The definition of general intelligence should not require perfection.
The Paper's Real Significance: Redefining the Agenda, Not Declaring Victory
If one accepts the "AGI is here" premise, the agenda fundamentally shifts:
| Old Agenda (AGI is future) | New Agenda (AGI is present) | |---|---| | Pursue breakthroughs to reach AGI | Manage risks and impacts of existing AGI systems | | Research "how to build AGI" | Research "how to coexist with AGI" | | Ethics is forward-looking | Ethics is urgent, real-time policy | | Governance frameworks can wait | Governance must catch up immediately |
The conclusion is clear: whether or not you accept the "AGI" label, the capability level of current systems forces us to think about risk, governance, and coexistence with entirely new frameworks.
Controversy and Reception
The comment has sparked ongoing academic debate:
An interesting data point: this Nature comment garnered 18 citations within 4 months, showing it has genuinely triggered scholarly discussion.
Conclusion: The Label Matters Less Than Action
The most valuable part of the article is its pragmatic pivot: whether or not you use the word "AGI," the capability reality of current systems has changed everything. Rather than debating labels, the priorities are:
1. Build governance frameworks matching system capabilities 2. Redesign education and employment systems 3. Ensure equitable distribution of the technology's benefits 4. Prevent misuse or concentration of capabilities
As the authors put it: "Eyes unclouded by anxiety can see the evidence clearly."