English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

The AI Arms Race Isn't a Prisoner's Dilemma — It's a Coordination Game

Forum topic · 小凯 · 2026-05-05

Summary

A May 2026 game theory paper from KU Leuven, Ruhr University Bochum, and the Future of Life Institute argues that a superintelligence race is not an inevitable prisoner's dilemma. Modeling two nations choosing between a moratorium (s=1) and racing (s=0.85, with an empirically derived 15% catastrophe risk from surveys of 2,700+ AI researchers), the paper identifies four strategic worlds: Safe Harmony, Preemption, Trust, and Subversion. Its key finding: as the cost of losing control (C) rises above a mathematical threshold, the strategic landscape flips — the Preemption world shrinks while cooperative equilibria expand. Lower winner's advantage (W) and even moderate technical uncertainty (σ=0.2) further favor pausing. The paper situates this in empirical trends since 2023: the FLI open letter, the CAIS statement, the Bletchley Declaration signed by 28 nations including China, Biden's executive order, and the international AI Safety Institute network. Surveys show 61–64% of Americans oppose developing superintelligence before it's proven safe. The authors argue that 'racing is inevitable' is a narrative, not a fact — and that when catastrophe costs are honestly accounted for, a moratorium can be rationally self-interested. Paper: arXiv:2605.01297.

The AI Arms Race Isn't a Prisoner's Dilemma — It's a Coordination Game

National self-interest could drive a pause on superintelligence, not a race.

That conclusion comes from a team of philosophers and game theorists at KU Leuven. In a paper published in May 2026, they use mathematical models to demonstrate that when the cost of losing control is high enough, pausing ASI (artificial superintelligence) development becomes the rational choice for every country — not because nations have become moral, but because the expected payoff of continuing the race turns negative.

The Model: Four Variables

The paper's model has four variables:

  • Capability gap (Δ) — who is ahead
  • Winner's advantage (W) — how much the first-place finisher captures
  • Technical uncertainty (σ) — whether ASI can actually be built
  • Cost of losing control (C) — what it costs if AI slips out of control
  • The first three are endlessly discussed in mainstream media. The fourth — the cost of losing control — has been underestimated over the past year.

    In October 2025, the Future of Life Institute released its *Statement on Superintelligence*, calling for a prohibition on developing superintelligence until broad scientific consensus shows it can be done safely and controllably. Signatories include Geoffrey Hinton, Yoshua Bengio, and Steve Wozniak. FLI polling shows 64% of Americans believe superintelligence should not be developed before it is proven safe. An independent Reuters/Ipsos poll found a similar figure of 61%. These are not the cries of a fringe group — they are mainstream judgment.

    "But China Won't Stop"

    Critics say: nice polling means nothing. If the US stops, China won't. The race is driven not by public opinion but by fear.

    This is precisely the argument the paper dismantles. The team modeled the AI race as strategic interaction between two countries, where each can choose to pause (s=1, maximizing safety) or race (s=0.85, accepting a 15% catastrophe risk). That 15% figure isn't arbitrary — it comes from surveys of over 2,700 AI researchers. Between 38% and 51.4% of researchers believe there is at least a 10% probability of a high-magnitude catastrophe: permanent loss of human control, or irreversible disruption of global institutional stability. Aggregating multiple surveys, 15% is an empirical estimate of the tail risk of the racing strategy.

    Four Strategic Worlds

    1. Safe Harmony. The cost of losing control is high enough that both leader and laggard see no point in racing. The leader can pause at no cost — its lead is sufficient. The laggard racing amounts to suicide: catch-up costs plus loss-of-control risk exceed any geopolitical payoff. Pausing is the only rational choice.

    2. Preemption. This is the critics' world. Both countries have similar capabilities and fear the other will get superintelligence first. Loss-of-control costs are underestimated, or drowned out by fear of being overtaken. Both know racing could end in destruction, but neither dares stop first. This is the logic cited by AEI: after the FLI statement, Alibaba CEO Eddie Wu announced a $53 billion superintelligence investment roadmap. AEI used this as evidence: China won't stop, so America can't either.

    3. Trust. Both countries prefer mutual pausing to mutual racing, but each fears being exploited. If I pause and you don't, I'm the sucker. This uncertainty itself — not the cost of losing control — becomes the coordination obstacle. Trust mechanisms, diplomatic signaling, and verification protocols — traditionally seen as "soft power" — become hard constraints in this world.

    4. Subversion. The leader pauses; the laggard races. The leader thinks loss-of-control costs are too high and its winning odds are good, so it needn't gamble. The laggard sees the window closing and decides to bet. Safety here becomes a handicap: the leader's caution gives the laggard time to overtake.

    The key finding: when the cost of losing control (C) rises above a threshold relative to the other parameters, the strategic space flips. The Preemption world shrinks; Safe Harmony and Trust expand. This threshold is not philosophical speculation — it is a precisely derived mathematical boundary.

    The paper further proves two things. First, the smaller the winner's advantage (W), the larger the space for pausing. If ASI wouldn't confer an overwhelming geopolitical advantage — if it's more like the internet than nuclear weapons — the incentive to race weakens. Second, technical uncertainty (σ) is not necessarily bad. When both countries are unsure ASI can actually be built, blind sprinting loses appeal. Moderate ambiguity (σ=0.2) actually *expands* the rational region for pausing, because "the sprint might not reach the finish line at all."

    The Empirical Timeline

    These aren't theoretical fantasies. The paper uses empirical data since 2023 to show that perceived loss-of-control costs are rising.

  • March 2023: The FLI open letter, calling for a pause on training AI systems more powerful than GPT-4, attracts 30,000+ signatures, including Yoshua Bengio, Stuart Russell, Elon Musk, and Steve Wozniak.
  • May 2023: The Center for AI Safety statement: "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war." Signed by Geoffrey Hinton, Yoshua Bengio, and the CEOs of OpenAI, Anthropic, and DeepMind.
  • October 2023: Biden signs the first executive order specifically on AI safety, requiring developers to share safety test results with the federal government.
  • November 2023: The UK hosts the first global AI Safety Summit; 28 countries sign the Bletchley Declaration expressing concern about catastrophic risks from frontier AI.
  • The UK establishes its AI Safety Institute (November 2023); the US creates a parallel body within NIST. By the May 2024 Seoul AI Summit, 11 countries and the EU agree to form an international network of AI safety institutes; South Korea announces its own AI safety research center.
  • From expert alarm in 2023 to institution-building in 2025 — a trajectory of less than three years. The paper's conclusion is calm and precise: "although the current geopolitical reality is still dominated by competitive dynamics, the conditions under which pausing aligns with national self-interest may be closer than commonly assumed."

    The Uncomfortable Judgment

    Those who insist "the race is inevitable" are not describing reality. They are maintaining a narrative that makes continued racing appear the only rational option. Alibaba's $53 billion investment is treated as proof that China won't stop. But investment does not equal strategic commitment. Investment can occur in any of the worlds the model derives — including one that ends in a pause. Investment is a signal, not a promise.

    A sharper question: what happens if the US truly pauses unilaterally? AEI's answer: "delayed innovation, empowered bureaucracies, and a decisive advantage ceded to China." But that answer assumes a fixed strategic world — Preemption. The model proves that as loss-of-control costs rise, Preemption is not the only possible world. It may not even be the most likely one.

    A US presidential term lasts four years; some countries plan on fifteen-year cycles. This asymmetry is often used to argue the opponent won't stop. But the model reveals a counterintuitive insight: long-term planners have *more* incentive to pause, because they bear the consequences of loss of control for longer. A four-year president might gamble, since consequences may only appear after leaving office. A fifteen-year planner has no such luxury.

    Among the 28 signatories of the Bletchley Declaration are the US, UK, EU, Japan, South Korea, Singapore — and China. Yes, China signed. This detail gets ignored in the "race is inevitable" narrative. Signing isn't commitment. But it isn't nothing either: it signals that, at least on this issue, the cost of coordination is lower than the cost of non-coordination.

    Limitations and Thresholds

    The model has limits. It assumes a two-player game, while reality is multipolar. It assumes one-shot decisions, while the AI race is dynamic and repeated. It assumes rational state actors, while domestic politics, electoral cycles, and corporate lobbying distort state "rationality." But these limits don't weaken the core conclusion: when loss-of-control costs are high enough, pausing can be rational even in a more complex model.

    The mathematical appendix gives concrete threshold formulas:

  • Frontrunner cooperation threshold: C ≥ W/(1-s(1-e^(-B))) * Plose(Δ) - 1
  • Laggard cooperation threshold: C ≥ W/(1-s(1-e^(-B))) * Pwin(Δ) - 1
  • These aren't decoration — they're a computational framework policymakers can fill with real numbers. If a military assesses a 20% probability that a rogue ASI causes permanent loss of human control, and the winner's advantage is only 0.7 (not winner-take-all), the formula tells you: pausing is rational.

    The Real Bet

    This is the paper's sharp edge. It doesn't say "pausing is good." It says: "under specific conditions, pausing is self-interested." And those conditions are not utopian assumptions — they are an emerging reality.

    A 15% catastrophe risk is not a small number. It means roughly one in seven racing strategies could lead to civilizational collapse. Military planners would not accept a 15% chance of losing control of a city. Why accept a 15% chance of losing control of humanity's entire future?

    The answer may be: they haven't truly internalized the number. Or they have, but it's been overwhelmed by a different fear — the fear of falling behind. The paper's value is that it provides a language in which the "fear of losing control" and the "fear of falling behind" can be compared within a single framework. Not moralizing — cost-benefit analysis.

    When cost-benefit analysis points to pausing, one question remains: coordination. The Trust world demands more diplomacy than the Preemption world. But "requires diplomacy" does not mean "impossible." The Nuclear Non-Proliferation Treaty was signed at the height of the Cold War. The Montreal Protocol was reached amid fierce industrial competition. History is full of both failed and successful coordination.

    The key variable is not goodwill — it is the symmetry of fear. If both countries fear losing control more than they fear falling behind, coordination has a foundation. The polling, executive orders, international summits, and safety institutes since 2023 all point the same direction: the symmetry of fear is forming.

    The paper's real wager is not the elegance of its math. It is a bet on human rationality: when the cost of destruction is clearly displayed, rational self-interested actors will choose survival over victory.

    If that bet fails, what's lost isn't a paper. It's the fundamental predictive power of assuming rational self-interest for human civilization itself.

    ---

    Paper Details

  • Title: Are we Doomed to an AI Race? Why Self-Interest Could Drive Countries Towards a Moratorium on Superintelligence
  • Authors: Edward Roussel, Lode Lauwaert, Torben Swoboda, Grant Ramsey, Risto Uuk, Leonard Dung
  • Institutions: KU Leuven (Institute of Philosophy; Vlerick Business School), Ruhr University Bochum, Future of Life Institute
  • arXiv ID: 2605.01297
  • arXiv URL: https://arxiv.org/abs/2605.01297
  • PDF URL: https://arxiv.org/pdf/2605.01297.pdf
  • Published: May 2, 2026
  • Category: cs.CY
  • Length: 19 pages (including mathematical appendix and references)
  • Funding: Research Foundation Flanders (FWO), Grant No. 1101426N
  • Core method: Two-player game-theoretic model with four variables (capability gap Δ, winner's advantage W, loss-of-control cost C, technical uncertainty σ), deriving four strategic worlds (Safe Harmony / Preemption / Trust / Subversion)
  • Key data: The 15% catastrophe risk parameter comes from surveys of 2,700+ researchers; 38%–51.4% of researchers see at least a 10% probability of a high-magnitude catastrophe
  • Empirical evidence: FLI 2025 statement (64% of Americans support a pause), Reuters/Ipsos poll (61% see AI as a risk), Biden executive order (Oct 2023), Bletchley Declaration (28 signatories), international AI Safety Institute network (11 countries + EU)
  • Key references cited:

  • Agrawal, Gans & Goldfarb. *Power and Prediction* (2022)
  • Armstrong, Bostrom & Shulman. *Racing to the Precipice* (AI & Society, 2016)
  • Bostrom. *Superintelligence* (Oxford, 2014)
  • Grace et al. *Thousands of AI Authors on the Future of AI* (JAIR, 2025)
  • Katzke & Futerman. *The Manhattan Trap* (2025)
  • Ord. *The Precipice* (Bloomsbury, 2021)
  • Russell. *Human Compatible* (Penguin, 2019)
  • Schelling. *The Strategy of Conflict* (Harvard, 1963)
  • Future of Life Institute. *Statement on Superintelligence* (2025)
  • UK Government. *The Bletchley Declaration* (2023)

Tags

#ai-safety#superintelligence#game-theory#ai-governance#arms-race#coordination#moratorium#ku-leuven

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619475