Paper Overview
Research Area: ML Authors: Derin Gezgin, Jim O'Connor, Tanner Goodwin Released: 2026-08-11 arXiv: 2508.03798
Summary
The authors introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of *Dark Souls: Remastered* as game-playing agent benchmarks through a Gymnasium-style interface. DSLE combines real-time combat, high-dimensional visual input, and sparse terminal rewards, with each environment step being a real action executed against the running game.
To support controlled comparison, they define DSLE-5, a representative five-boss subset spanning a melee fight, a spatially constrained arena, an environmental-hazard fight, a multi-target fight, and a fast final-boss fight, recommended as the starting suite for agents built on DSLE.
On DSLE-5, the paper evaluates a random policy, an expert system, an evolutionary baseline, and PPO and DQN agents trained from visual input.
Key Findings
- The expert system and evolutionary baseline each defeated the game's tutorial boss, the Asylum Demon, with peak win rates of 63% and 43% respectively.
- None of the five methods defeated the other four bosses in DSLE-5.
- PPO and DQN showed no measurable learning within a budget that already required tens of hours of wall-clock time per run — at most 0.33% win rate on the tutorial boss and 0% elsewhere.
- A broader study ran the evolutionary baseline across all 22 encounters; even with all 50 levels of stats maxed, it won only a handful of additional early-game bosses and nothing beyond.
- Failure cases ranged from deaths in under 10 seconds in narrow, multi-target encounters to stalemates lasting nearly a minute with almost no damage dealt. The paper reports these via survival time and damage dealt rather than win rate alone.
Original Abstract (excerpt)
> We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of Dark Souls: Remastered as game-playing agent benchmarks through a Gymnasium-style interface. DSLE combines real-time combat, high-dimensional visual input, and sparse terminal rewards, with each environment step being a real action executed against the running game. To support controlled comparison, we define DSLE-5, a representative five-boss subset, spanning a melee fight, a spatially constrained arena, an environmental-hazard fight, a multi-target fight, and a fast final-boss fight, that we recommend as the starting suite for agents built on DSLE. On DSLE-5 we evaluate a random policy, an expert system, an evolutionary baseline, and PPO and DQN agents trained from vi...
Link: arXiv:2508.03798
---
*Auto-collected on 2026-08-12*