Overview
Field: NLP Authors: Shashwat Goel, Nikhil Chandak, Arvindh Arun Published: 2026-05-16 arXiv: 2505.08630
Abstract
AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently measure this capability for realistic use-cases, the authors propose building grounded simulations that replay real-world events in the order they occurred.
FutureSim is built on this idea: agents forecast world events beyond their knowledge cutoff while interacting with a chronological replay of the world — real news articles arriving and questions resolving over the simulated period. Frontier agents are evaluated in their native harness, testing their ability to predict world events over a three-month period from January to March 2026.
Key Findings
- FutureSim reveals a clear separation in agent capabilities: the best agent's accuracy is only 25%, and many agents have worse Brier skill scores than making no prediction at all.
- Careful ablations show the benchmark enables realistic study of emerging research directions: long-horizon test-time adaptation, search, memory, and reasoning under uncertainty.
Conclusion
The authors hope this benchmark design paves the way for measuring AI's progress on open-ended adaptation over long time horizons in the real world.
--- *Auto-collected on 2026-05-16*