Paper Overview
- Field: Machine Learning
- Authors: Elias Fernández Domingos, The Anh Han
- Posted: 2026-07-28
- arXiv: 2607.26034
- Data did not support the pre-registered comparisons between risk levels, nor a role for elicited risk preferences.
- Exploratory analyses driven by the repeated structure of the task suggest unsafe behaviour is shaped less by risk preference and more by the evolving strategic state of the race:
- Participants were more likely to choose unsafe development after their opponent did so.
- Leading decreased unsafe play; falling behind increased unsafe play.
- First-round choices predicted subsequent behaviour (early behavioural inertia).
Summary
Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development.
The study uses a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10%, 60%, or 90%; the race's competitive structure was held constant, and only this maximum risk varied.
Key Findings
Evolutionary Model
To explain these effects, the authors introduce a simplified evolutionary model with four strategies:
1. Always Safe 2. Always Unsafe 3. Conditionally Safe 4. Conditionally Antisocial-Safe
The model reproduces the treatment effects and demonstrates how conditional unsafe behaviour is favoured by competitive race dynamics.
Implications
Together, the experiment and model indicate that unsafe development may arise from early behavioural inertia, opponent behaviour, and the fear of falling behind—rather than from risk appetite alone. This suggests AI governance policy should focus on reducing competitive pressures and fostering cooperation in AI development, rather than solely addressing individual risk preferences.
--- *Auto-collected 2026-07-30*