Paper Overview
Field: Machine Learning Authors: Lennon J. Shikhman, Michael Galarnyk, Aadi Dash, Nicholas A. Welsh arXiv: 2607.27188
Abstract
Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation, while a chronological NIFTY benchmark tests only held-out market prices. A two-component lognormal mixture has the lowest aggregate price, L^1, Wasserstein, and fixed-tail errors on the synthetic benchmark. Learned operators retain narrower strengths: DeepONet reduces 1% quantile and variance error by 39.0% and 34.6% relative to the mixture, and a quote transformer reduces L^1 by 16.4% on the structurally misspecified Merton family. A numerical conditioning analysis explains why these rankings can differ: after enforcing mass and forward constraints, 95 of 126 pricing directions are numerically zero, and two densities with L^1 distance 0.061 produce identical prices on covered strikes.
On 524 held-out NIFTY calls, validation-selected test-time adaptation lowers DeepONet RMSE by 28.3%, but per-maturity mixture and SVI fits remain considerably more accurate.
Key Findings
- Pricing accuracy ≠ density recovery: the paper formally separates market-price performance from latent density recovery quality.
- Two-component lognormal mixture wins on aggregate price, L^1, Wasserstein, and fixed-tail errors in the synthetic benchmark.
- DeepONet improves 1% quantile error by 39.0% and variance error by 34.6% over the mixture.
- Quote transformer reduces L^1 error by 16.4% on the misspecified Merton family.
- Ill-conditioning: after mass and forward constraints, 95 of 126 pricing directions are numerically zero; two densities with L^1 distance 0.061 yield identical prices on covered strikes.
- NIFTY benchmark: test-time adaptation cuts DeepONet RMSE by 28.3% on 524 held-out calls, yet per-maturity mixture and SVI fits remain more accurate.