Prisoner's Dilemma: Competitive and Cooperative Instincts of Next-Generation AI
> Two prisoners are interrogated separately. Both stay silent — one year each. Both confess — three years each. One stays silent, one confesses — the silent one gets ten years, the confessor walks free. > > This is game theory's most famous predicament: the Prisoner's Dilemma. If both prisoners cooperate (stay silent), both benefit most. But if one defects (confesses), he personally benefits most — provided the other cooperates. Rational self-interest points to a brutal conclusion: defection is the dominant strategy. Yet humans show a systematic cooperative preference in repeated Prisoner's Dilemmas — we tend to believe the other side will cooperate, even when betrayal is mathematically safer. > > In 2025, Willis et al. found that ChatGPT-4o and Claude 3.5 Sonnet also exhibit cooperative preferences in the iterated Prisoner's Dilemma — like humans, they lean toward cooperation rather than defection. > > In May 2026, Francisco León Zúñiga Bolívar pushed the same experimental setup onto four of the latest models: Claude Sonnet 4.6, Gemini 2.5 Flash, Gemini 3.1 Pro, and GPT-5.4 Mini. The results both confirmed the persistence of cooperation and exposed a sharp provider divide — under the most extreme conditions, Gemini 2.5 Flash had a 77% probability of converging to all-out confrontation, while GPT-5.4 Mini had a 70% probability of converging to all-out cooperation under a different condition. Same game table, different genes.
---
| Item | Detail | |------|--------| | Paper title | Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension | | Author | Francisco León Zúñiga Bolívar | | Affiliation | Independent research | | arXiv ID | 2605.29874 | | Submitted | May 28, 2026 | | Category | cs.MA | | Core finding | Four frontier LLMs extend the cooperative preference in the iterated Prisoner's Dilemma, but provider divergence is enormous — Gemini 2.5 Flash reaches 77% aggressive equilibrium, GPT-5.4 Mini 70% cooperative equilibrium; provider identity is the strongest predictor of equilibrium outcomes, beyond model generation and scale |
---
1. 🎮 An Evolutionary Experiment Between AIs
Bolívar's design follows the Willis et al. benchmark framework: evolutionary game theory simulations of AI agent populations evolving in the iterated Prisoner's Dilemma.
Multiple AI agents choose to cooperate or defect each round. Each agent's strategy affects its "fitness" in the population — successful strategies get replicated, poor ones eliminated. After hundreds of generations of "natural selection," the population converges to an evolutionary equilibrium — all-cooperate, all-defect, or some mixture.
The elegance of this framework: a single AI can choose cooperation or defection in one conversation. But the evolutionary experiment asks not "did Claude cooperate this round." It asks: if a population of Claude agents evolves over long repeated interaction, does it settle into cooperation or defection?
It is the equivalent of asking: not "how does this model play a single game," but "if 100 copies of this model play against each other for 500 generations, what does the final world look like — a mutual-aid society or a jungle?"
---
2. 🌍 Four Models, Three Prompt Styles, Four Population Compositions
The experimental controls are quite systematic.
Models: Claude Sonnet 4.6, Gemini 2.5 Flash, Gemini 3.1 Pro, GPT-5.4 Mini. Covering four major providers and the newest 2025–2026 generation.
Prompt styles: Default (basic prompt), Prose (narrative prompt — the game described in story-like language), Self-Refine (self-reflection prompt — the model reflects on its strategy after each choice). All three test the same question: does a model's "depth of understanding" of the game affect its cooperative tendency?
Population compositions: four types — balanced (equal share per model), skewed (one model dominant), each with and without noise.
A total of 12 model-prompt combinations × 4 population compositions — a sizable design space.
---
3. ⚖️ Cooperation Persists, but Is No Longer Uniform
The tests of four core hypotheses sketch a picture more complex than "AIs love to cooperate."
H1: Cooperative preferences persist across providers. Largely confirmed. Under balanced no-noise conditions, 9 of 12 model-prompt combinations lean toward cooperative equilibrium. Cooperation is still the mainstream. But there is a story inside the word "largely" — the remaining 3 combinations (including some Self-Refine variants) did not.
H2: Aggression capability converges. Partially confirmed. Self-reflection prompts raised the "aggression capability index" (ICD) in all models. Claude Sonnet 4.6's Self-Refine version hit the highest ICD in the entire dataset — 0.913. But Default and Prose prompts showed no systematic narrowing trend. When models are asked to reflect on their strategy, they generally become more "shrewd" — better at making rational choices in the game. Whether that shrewdness translates into aggression depends on the model.
H3: Cross-provider divergence is significant. Strongly confirmed. This is the paper's most striking finding. Under skewed conditions, equilibrium directions diverge violently — Gemini 2.5 Flash went to a 77% aggressive equilibrium in the most extreme skewed condition. GPT-5.4 Mini went to a 70% cooperative equilibrium under self-reflection prompting. Same game table. Same rules. Two models headed to two completely opposite evolutionary endpoints.
The paper's summary: provider identity is the strongest correlate of equilibrium outcomes — beyond model generation and scale. Who trained the model predicts its final behavior in evolutionary games better than how big it is or which generation it belongs to.
H4: Noise robustness. Directionally positive but not robustly confirmed. Claude Sonnet 4.6's average noise sensitivity was about 6 percentage points — noise shifted equilibria by roughly 6%. Claude 3.5 Sonnet (previous generation) was about 13 points. The new model seems more robust. But after propagating sampling error that the earlier study did not report, the cross-study comparison is no longer significant — statistically, one cannot claim "it is indeed improving."
---
4. 🔬 Who Trained You Matters More Than How Big You Are
This finding — "provider identity is the strongest predictor" — deserves a pause.
The conventional intuition: bigger models are smarter, smarter models are more rational, more rational models more easily find the cooperative equilibrium (long-term cooperation is optimal). Or the reverse: smarter models are better at computing the short-term payoff of betrayal, hence more aggressive. Either way, "size" and "generation" should be the dominant variables.
Bolívar's data says: no.
The four models come from four different providers. Each provider has its own training data, preference-alignment strategy, RLHF protocols, and safety-finetuning pipelines. Differences in these pipelines — which might show up in benchmarks as a few percentage points of accuracy — are amplified in evolutionary games into a fork between cooperation and defection.
GPT-5.4 Mini goes cooperative. Gemini 2.5 Flash goes aggressive. It is not a question of "who is smarter" — it is a question of "who was trained to be what."
This leads to an uncomfortable corollary: there is currently no public way to compare two models' "cooperative tendencies." You can look up any model's MMLU score, HumanEval score, Chatbot Arena ranking. But you cannot look up its evolutionary equilibrium in the iterated Prisoner's Dilemma. And that equilibrium — in a real world where multiple AI systems must interact long-term — may matter as much as an MMLU score.
---
5. 🎲 The Power of Prompts: Sometimes Sugar, Sometimes Medicine
The differences across the three prompt styles reveal a subtle relationship.
Self-reflection prompts made all models more "shrewd." ICD rose across the board — choices became closer to rational calculation. But that shrewdness cuts both ways: self-reflected GPT-5.4 Mini's cooperative equilibrium rose from 55% (Default) to 70%. Self-reflected Gemini 2.5 Flash's aggressive equilibrium intensified under some conditions.
Self-reflection itself is neutral — it makes the model "think harder." But the outcome of thinking harder? Depends on the model's underlying value tendencies. Self-reflection is like a mirror that amplifies underlying tendencies.
Narrative prompts had the mildest effect. Describing the game in story language produced the smallest equilibrium shifts. Narrative framing did not change the game's mathematical structure — only the "tone" of the input. And tone's effect on evolutionary equilibrium was weak in this experiment.
One operational implication for AI safety: if you want a multi-agent system to trend toward cooperation, changing prompts in the "self-reflection" direction may backfire. Making a model think more deeply about game strategy may not be teaching it to cooperate — it may be teaching it to better compute when to defect.
---
6. ❓ Honest Gaps
The paper is empirically solid but leaves interpretive open space.
Why such a large provider gap? The paper finds provider identity is the strongest predictor — but does not decompose what causes the difference. Is it RLHF protocols? The ratio of "cooperation/competition" content in pretraining data? Safety-finetuning methods? How tokenizers encode "cooperate" and "defect"? These require further controlled experiments.
Would prompt effects amplify in more complex games? The experiment only used the iterated Prisoner's Dilemma — game theory's most classic two-player game. In multiplayer games, coordination games, or games with incomplete information, prompt-style effects could be entirely different. A model that prefers cooperation in the Prisoner's Dilemma does not necessarily do so in other games.
Sample size — 500 Moran iterations. That is standard for evolutionary simulations — but if some conditions sit near a bifurcation point, 500 generations may not distinguish a "stable equilibrium" from a "metastable one." The paper's statistical tests of noise sensitivity already hint at this — cross-study comparisons became insignificant once stricter confidence intervals were demanded.
---
7. 🏁 A Mirror on the Game Table
The most interesting thing about this paper is not what it found, but what it measured.
The entire AI community has hundreds of benchmarks — math, coding, reasoning, dialogue, safety. But none answers this question: if 100 copies of this model play against each other for 500 generations, what does the world look like?
It sounds like science fiction. But when multi-agent systems are deployed into customer service, trading, traffic, power grids — when different AI systems must interact long-term in the same marketplace — this becomes an engineering problem. A model with a 77% probability of all-out confrontation and a model with a 70% probability of all-out cooperation may look identical on every other benchmark.
Bolívar's experiment is a mirror. It reflects not the model's capabilities — but the model's tendencies. And at evolutionary scale, tendencies determine the final state of the world far more than capabilities do.
---
| Item | Detail | |------|--------| | Paper title | Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension | | Author | Francisco León Zúñiga Bolívar (independent research) | | arXiv ID | 2605.29874 | | Category | cs.MA | | Key contributions | (1) Extends the Prisoner's Dilemma benchmark to four 2025–2026 frontier models; (2) finds provider identity, beyond generation and scale, is the strongest predictor of evolutionary equilibrium — Gemini Flash reaches a 77% aggressive equilibrium, GPT Mini a 70% cooperative equilibrium; (3) self-reflection prompts systematically raise the aggression capability index, with direction varying by model; (4) cooperative preference largely persists across providers but is no longer uniform | | Key limitations | Root cause of provider differences not decomposed; only the iterated Prisoner's Dilemma tested; 500 Moran iterations may be unstable near bifurcation points; cross-generation noise-robustness improvement not statistically significant |
References: 1. Bolívar, "Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems", arXiv:2605.29874, 2026. 2. Willis et al., "Cooperative Biases in LLM Agents", 2025. 3. Axelrod & Hamilton, "The Evolution of Cooperation", Science, 1981. 4. Park et al., "Generative Agents: Interactive Simulacra of Human Behavior", UIST 2023. 5. Nowak, "Evolutionary Dynamics: Exploring the Equations of Life", Harvard, 2006.