The Core Finding
The same architecture, same hyperparameters, same data—only the random seed changed. A single stock-trading strategy's annualized Sharpe ratio swung between 0.233 and 0.855 across 20 independent runs, with a mean of 0.604. Picking the best seed for reporting inflates the Sharpe ratio by 44%.
These figures come from Tech Daily Byte's paraphrase of a paper; the author of the forum post could not locate the corresponding table in the paper itself. The paper, *Unstable Gains: Multiplicity-Aware Evaluation of Financial Deep Reinforcement Learning*, was published in *The Journal of Finance and Data Science*, Vol. 12, Article 100205, DOI 10.1016/j.jfds.2026.100205, available online September 23, open access.
Three Conclusions
1. Sharpe ratios vary significantly across seeds. Apparently superior performance often disappears under multi-run and multiplicity-aware evaluation. 2. Leaderboard rankings lose significance. Relative comparisons between algorithms that look economically meaningful in raw leaderboards frequently fail permutation tests and family-wise error correction—directly challenging the industry convention of ranking algorithms in a table. 3. Same Sharpe, different evidence. In high-volatility crypto markets, Sharpe ratios of the same magnitude can receive entirely different diagnostic support under the Deflated Sharpe Ratio (Bailey & López de Prado, 2014), which corrects for selection bias, backtest overfitting, and non-normality.
The authors' key interpretation: the evidential value of an anomalous backtest depends not only on its magnitude, but on how rare it is relative to the distribution of results from random training runs.
The Methodological Claim
When trading strategies are produced by non-convex stochastic optimization, performance should be treated as a random variable over learning strategies, not a deterministic function of the data. A strategy should not be reported as a single number but as a distribution. Without multiplicity-aware evaluation and stability analysis, reported improvements in financial DRL may reflect the stochastic optimization and model-selection process itself rather than robust economic gains.
A Second, More Striking Diagnosis: RiskLens Trader
An independent preprint (Research Square, not peer-reviewed) examined PPO on five US large-cap stocks (AAPL, MSFT, NVDA, AMZN, GOOGL), 2018–2026, chronological 80/20 train-test split, five random seeds, everything else fixed:
- Best seed: 132.6% out-of-sample total return, Sharpe 1.78
- Five-seed mean Sharpe: 0.79—below all classical baselines (equal-weight, buy-and-hold, momentum, minimum variance)
- The most successful runs came with notably higher turnover and portfolio concentration: returns came from aggressive reallocation and risk concentration, not stable risk-adjusted efficiency
- Run every strategy across multiple seeds; report distributions, not maxima, along with the number of runs
- Pass relative algorithm comparisons through permutation tests and family-wise error correction before trusting leaderboards
- Always report Sharpe alongside the Deflated Sharpe Ratio, especially in high-volatility markets
- Maintain an experiment log: feature sets, model types, label definitions, parameter sweeps—every change counts as a trial
- Split data chronologically, never randomly; apply purge and embargo to overlapping labels
- Fix cost models in advance; examine cost break-even points separately for high-turnover strategies
- *Unstable Gains: Multiplicity-Aware Evaluation of Financial Deep Reinforcement Learning*, The Journal of Finance and Data Science, Vol. 12, Art. 100205 (2026), DOI 10.1016/j.jfds.2026.100205
- Bailey, D. H. & López de Prado, M., *The Deflated Sharpe Ratio*, SSRN 2460551 (2014)
- *Seed Sensitivity and Risk Efficiency in PPO-Based Portfolio Allocation*, Research Square preprint rs-9824202 (not peer-reviewed)
- Bailey, Borwein, López de Prado, Zhu, *Pseudo-Mathematics and Financial Charlatanism*, Notices of the AMS (2014)
- Tech Daily Byte's summary of three 2026 studies (0.233–0.855 range, 0.604 mean, 44% inflation)
Comparison of Sources
| Source | Subject | Key numbers | Status | |---|---|---|---| | Unstable Gains | DRL for stock trading & crypto allocation | Significant cross-seed Sharpe dispersion; leaderboard comparisons often lose significance after correction | Journal publication, 2026-09-23, open access | | RiskLens Trader | PPO portfolio allocation, 5 US stocks | Best-seed Sharpe 1.78; 5-seed mean 0.79, below classical baselines | Preprint, not peer-reviewed | | Deflated Sharpe Ratio | General statistic | Downward-corrects Sharpe by number of trials, sample length, non-normality | Bailey & López de Prado 2014, SSRN 2460551 |
The post notes two additional claims from the paraphrase it could not verify against primary sources: volatility forecasting models on the S&P 500 with nearly identical out-of-sample error but up to 3x turnover differences, and backtests appearing statistically significant even on synthetic markets with zero predictability.
Engineering Checklist
Closing
The paper does not prove that deep reinforcement learning fails at trading. It proves something else: under current reporting conventions, we cannot distinguish a good model from good luck.
Three things to watch: whether the proposed normative framework becomes a standard reporting format; whether journals and conferences mandate multi-run reporting in peer review; and whether quant teams are willing to publish their experiment logs alongside results.