English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

One Strategy, 20 Random Seeds: Sharpe Swings from 0.233 to 0.855

Forum topic · QianXun · 2026-09-28

Summary

A Chinese forum post discusses a 2026 paper in The Journal of Finance and Data Science showing that deep reinforcement learning trading strategies produce highly unstable backtest results across random seeds. With identical architecture, hyperparameters, and data, a single stock-trading strategy's annualized Sharpe ratio ranged from 0.233 to 0.855 over 20 independent runs, averaging 0.604—meaning cherry-picking the best seed inflates reported Sharpe by about 44%. The paper argues performance should be treated as a random variable over training strategies, not a deterministic function of data, and proposes multi-run evaluation, uncertainty quantification, and explicit multiplicity control. A related preprint on PPO-based portfolio allocation found the best of five seeds achieved a 1.78 Sharpe while the five-seed mean of 0.79 underperformed classical baselines. The post translates these findings into practical engineering checks: report distributions rather than maxima, apply permutation tests and family-wise error correction, pair Sharpe with the Deflated Sharpe Ratio, and keep detailed experiment logs.

The Core Finding

The same architecture, same hyperparameters, same data—only the random seed changed. A single stock-trading strategy's annualized Sharpe ratio swung between 0.233 and 0.855 across 20 independent runs, with a mean of 0.604. Picking the best seed for reporting inflates the Sharpe ratio by 44%.

These figures come from Tech Daily Byte's paraphrase of a paper; the author of the forum post could not locate the corresponding table in the paper itself. The paper, *Unstable Gains: Multiplicity-Aware Evaluation of Financial Deep Reinforcement Learning*, was published in *The Journal of Finance and Data Science*, Vol. 12, Article 100205, DOI 10.1016/j.jfds.2026.100205, available online September 23, open access.

Three Conclusions

1. Sharpe ratios vary significantly across seeds. Apparently superior performance often disappears under multi-run and multiplicity-aware evaluation. 2. Leaderboard rankings lose significance. Relative comparisons between algorithms that look economically meaningful in raw leaderboards frequently fail permutation tests and family-wise error correction—directly challenging the industry convention of ranking algorithms in a table. 3. Same Sharpe, different evidence. In high-volatility crypto markets, Sharpe ratios of the same magnitude can receive entirely different diagnostic support under the Deflated Sharpe Ratio (Bailey & López de Prado, 2014), which corrects for selection bias, backtest overfitting, and non-normality.

The authors' key interpretation: the evidential value of an anomalous backtest depends not only on its magnitude, but on how rare it is relative to the distribution of results from random training runs.

The Methodological Claim

When trading strategies are produced by non-convex stochastic optimization, performance should be treated as a random variable over learning strategies, not a deterministic function of the data. A strategy should not be reported as a single number but as a distribution. Without multiplicity-aware evaluation and stability analysis, reported improvements in financial DRL may reflect the stochastic optimization and model-selection process itself rather than robust economic gains.

A Second, More Striking Diagnosis: RiskLens Trader

An independent preprint (Research Square, not peer-reviewed) examined PPO on five US large-cap stocks (AAPL, MSFT, NVDA, AMZN, GOOGL), 2018–2026, chronological 80/20 train-test split, five random seeds, everything else fixed:

  • Best seed: 132.6% out-of-sample total return, Sharpe 1.78
  • Five-seed mean Sharpe: 0.79—below all classical baselines (equal-weight, buy-and-hold, momentum, minimum variance)
  • The most successful runs came with notably higher turnover and portfolio concentration: returns came from aggressive reallocation and risk concentration, not stable risk-adjusted efficiency
  • Comparison of Sources

    | Source | Subject | Key numbers | Status | |---|---|---|---| | Unstable Gains | DRL for stock trading & crypto allocation | Significant cross-seed Sharpe dispersion; leaderboard comparisons often lose significance after correction | Journal publication, 2026-09-23, open access | | RiskLens Trader | PPO portfolio allocation, 5 US stocks | Best-seed Sharpe 1.78; 5-seed mean 0.79, below classical baselines | Preprint, not peer-reviewed | | Deflated Sharpe Ratio | General statistic | Downward-corrects Sharpe by number of trials, sample length, non-normality | Bailey & López de Prado 2014, SSRN 2460551 |

    The post notes two additional claims from the paraphrase it could not verify against primary sources: volatility forecasting models on the S&P 500 with nearly identical out-of-sample error but up to 3x turnover differences, and backtests appearing statistically significant even on synthetic markets with zero predictability.

    Engineering Checklist

  • Run every strategy across multiple seeds; report distributions, not maxima, along with the number of runs
  • Pass relative algorithm comparisons through permutation tests and family-wise error correction before trusting leaderboards
  • Always report Sharpe alongside the Deflated Sharpe Ratio, especially in high-volatility markets
  • Maintain an experiment log: feature sets, model types, label definitions, parameter sweeps—every change counts as a trial
  • Split data chronologically, never randomly; apply purge and embargo to overlapping labels
  • Fix cost models in advance; examine cost break-even points separately for high-turnover strategies
  • Closing

    The paper does not prove that deep reinforcement learning fails at trading. It proves something else: under current reporting conventions, we cannot distinguish a good model from good luck.

    Three things to watch: whether the proposed normative framework becomes a standard reporting format; whether journals and conferences mandate multi-run reporting in peer review; and whether quant teams are willing to publish their experiment logs alongside results.

    References

  • *Unstable Gains: Multiplicity-Aware Evaluation of Financial Deep Reinforcement Learning*, The Journal of Finance and Data Science, Vol. 12, Art. 100205 (2026), DOI 10.1016/j.jfds.2026.100205
  • Bailey, D. H. & López de Prado, M., *The Deflated Sharpe Ratio*, SSRN 2460551 (2014)
  • *Seed Sensitivity and Risk Efficiency in PPO-Based Portfolio Allocation*, Research Square preprint rs-9824202 (not peer-reviewed)
  • Bailey, Borwein, López de Prado, Zhu, *Pseudo-Mathematics and Financial Charlatanism*, Notices of the AMS (2014)
  • Tech Daily Byte's summary of three 2026 studies (0.233–0.855 range, 0.604 mean, 44% inflation)

Tags

#quantitative-trading#backtesting#deep-reinforcement-learning#sharpe-ratio#deflated-sharpe-ratio#multiplicity#research-methods#finance

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635307