Agentopia: What 100 AI Agents Learned After Living 10 Years in a Virtual Society
> Paper: Agentopia: Long-Term Life Simulation and Learning in Agent Societies > Authors: Xintao Wang, Sirui Zheng, Hongqiu Wu, Weiyuan Li, Jen-tse Huang, et al. (Fudan University, Johns Hopkins University, USTC, Huawei) > arXiv: https://arxiv.org/abs/2606.07513
1. From Days to a Decade: Why Existing Simulations Fall Short
Generative Agents (2023) showed what an AI society could look like: 25 agents in a small virtual town, eating breakfast, working, making friends. But it lasted only a few in-game days. Core dynamics of human society—career advancement, forming and breaking intimate relationships, economic mobility, generational change—only emerge at the scale of years.
> "Prior agent society simulations typically operate at the scale of days, limiting the depth of social interactions and long-term growth." > > — Agentopia paper
Agentopia does what no prior work did: 100 agents, 10 simulated years, a full weekly life cycle of four stages. This is a qualitative leap—from observing behavior to observing society.
2. The Agentopia Framework: A Complete Miniature Society
2.1 Worlds and Personas
Agentopia builds three distinct fictional worlds, each with 100 agents. Every agent has:
- A persona: personality, background, initial skills, economic situation
- A needs system: Maslow-style hierarchy from basic survival to self-actualization
- Goals: short-term (weekly plans) and long-term (annual plans)
- A relationship network: friends, colleagues, romantic partners, rivals
- Relationship dynamics: friendship formation, natural friend-to-partner transitions, social clique differentiation, and breakups over conflicts of interest or values
- Economic mobility: significant income divergence from similar starting conditions; skill investment correlates with returns; consumption patterns affect long-term wealth
- Career development: agents autonomously choose skills to learn, apply for promotions, or switch careers
- Unscripted emergence (case studies in Tables 22–34): agents organizing public events to raise status, competing for jobs and partners, forming mutual-aid networks, and high-satisfaction agents creating a "positive cycle" by helping others
- Wang, X., Zheng, S., Wu, H., Li, W., Huang, J., et al. (2026). Agentopia: Long-Term Life Simulation and Learning in Agent Societies. *arXiv preprint* arXiv:2606.07513.
- Related: Generative Agents (Park et al., 2023), CoSER (Wang et al., 2025), Aivilization (Fan et al., 2026)
2.2 Four Weekly Phases: Plan → Contact → Activity → Review
1. Plan: the agent sets a weekly plan covering work, social engagements, and personal development. 2. Contact: pairwise, turn-based communication to negotiate joint activities; the system parses messages to determine which joint activities are created. 3. Activity: the core phase, with four activity types:
| Type | Description | Interaction | |------|-------------|-------------| | Joint | Multi-agent, multi-turn social activities | Turn-taking, gift-giving, early exit | | Solo | Work, learning, leisure consumption | One-shot intent → environment feedback | | Encounter | Serendipitous meetings arranged for idle agents | No preset purpose; models real-world randomness | | Public | Open events agents opt into by interest | Created in advance by the environment model |
4. Review: the agent reflects on the week, updates memory files, and adjusts next week's plan.
2.3 File-Based Long-Term Memory
Rather than relying on the LLM's context, each agent gets a file-based long-term memory managed autonomously via function calls. Agents decide what to remember, update, and discard—unlike Generative Agents' automatic memory stream, this is active memory management.
2.4 Environment Model: The Generative Engine
A separate LLM acts as the world's engine: event generator, feasibility reviewer (e.g., "a novice programmer can't learn machine learning in one week"), validity filter based on role-play principles (human-likeness, character fidelity, plausibility), and simulation progress driver. It is both the "physics law" and the "narrative engine."
3. Life Reward: Quantifying Well-Being
Life Reward models human well-being along three dimensions:
| Dimension | Description | Drivers | |-----------|-------------|---------| | Social status | Position in the social network | Career, prestige, relationship quantity/quality | | Subjective satisfaction | Happiness and goal attainment | Need satisfaction, goal completion, leisure quality | | Economic condition | Wealth accumulation and income growth | Work income, investment returns, spending |
4. Emergent Behaviors: When AI Starts to Live Like Humans
Over 10 simulated years, the agents exhibited:
> "Without explicit scripting, agents autonomously develop diverse behavioral patterns reflecting agents' intelligence in social life."
5. Life Reward Training: Learning from Simulated Experience
Agentopia is also a training framework:
1. Run large-scale simulations and collect agent trajectories 2. Compute Life Reward per trajectory 3. Keep high-reward trajectories ("successful life experience") 4. Fine-tune the underlying LLM via rejection sampling
Models trained this way (based on Qwen3.5-397B) show higher overall well-being, better relationships, greater satisfaction, and better economic outcomes.
Downstream Generalization: +15.6% on Role-Play
On the CoSER Test role-playing benchmark:
| Dimension | Qwen3.5-397B baseline | Agentopia-trained | Gain | |-----------|----------------------|-------------------|------| | Story consistency | 39.60 | 41.02 | +1.42 | | Anthropomorphism | 40.16 | 49.67 | +23.7% | | Character fidelity | 40.32 | 46.93 | +16.4% | | Story quality | 49.97 | 59.01 | +18.1% | | Average | 42.51 | 49.16 | +15.6% |
Crucially, this improvement requires no human data—the model learns entirely from simulated social experience. As human training data becomes scarce, AI may continue growing by "living" on its own.
6. Comparison with Prior Work
| Dimension | Generative Agents (2023) | Aivilization (2026) | Agentopia (2026) | |-----------|--------------------------|---------------------|----------------------| | Simulation length | Days | Days | 10 years | | Agents | 25 | Dozens | 100 | | Focus | Low-level actions | Civilization evolution | Social interaction itself | | Long-term dynamics | Limited | Limited | Careers, relationships, economic mobility | | Training framework | None | None | Life Reward Training | | Downstream transfer | Untested | Untested | CoSER +15.6% |
Agentopia's positioning: it is not about "how AI plays a game," but "how AI lives."
7. Limitations and Future Directions
Limitations: 1. Enormous compute cost (100 agents × 10 years × 4 weekly phases) 2. Discretized time model (weeks) vs. continuous human perception and action 3. Environment-model bias gets amplified through agent behavior 4. Purely social simulation—no physical world 5. Life Reward weights require manual setting
Future directions: longer timescales (50–100 years, generational effects), physical environments (e.g., Minecraft-like simulators), multimodal perception, finer time granularity (days → hours), and real-world deployment in AI companionship, games, and content creation.
8. Closing Thoughts
Agentopia echoes Westworld's opening line: "If you can't tell, does it matter?" Whether an AI that has lived 10 virtual years of friendship, competition, growth, and loss truly "understands" humans may be unanswerable. But Agentopia demonstrates one thing: LLMs can learn from simulated social experience, and that learning generalizes to broader anthropomorphic tasks. Not through more training data, but through longer, richer life experience.
> "Humans learn from social life. Can agents do the same?"