Paper Overview
- Field: Machine Learning (ML)
- Authors: Tomasz R. Bielecki, Thibaut Mastrolia, Haoze Yan
- Published: 2026-08-19
- arXiv: 2608.19151
- A finite-dimensional Markovianization procedure and algorithm that approximates multivariate Hawkes processes using mixtures of exponential kernels.
- A convergence proof showing that the Markovianized approximation, its intensity, and the value function converge to the original non-Markovian process and the value of the primal problem.
- A continuous-time deterministic policy gradient method, called Hawkes-CT DDPG, built on the Markovianized approximation.
- A model-free algorithm that solves the non-Markovian Hawkes-driven optimization using only observations of event times, SDE realizations, and a set of chosen decay filters, while the Hawkes kernel coefficients remain unknown.
- Empirical comparison of the Hawkes-CT DDPG approach against discrete-time reinforcement learning techniques across three kernel types: simple exponential, Erlang, and power-law kernels.
Summary
The paper studies stochastic control of multivariate Hawkes-driven stochastic differential equations (SDEs) using machine learning algorithms in a non-Markovian setting. Because the Hawkes intensity carries memory that depends on the entire event history, the control problem lies outside the classical stochastic control framework, except for special Markovian kernels.
Key contributions
Original abstract (excerpt)
> We study stochastic control of multivariate Hawkes-driven stochastic differential equations with machine learning algorithms in a non-Markovian setting. Due to the path dependence of the memory of the Hawkes intensity, this problem does not fall within classical stochastic control theory outside particular Markovian kernels. We first develop a finite-dimensional Markovianization procedure and algorithm to approximate multivariate Hawkes processes with mixtures of exponential kernels. We prove the convergence of the Markovianized approximation of the Hawkes process, its intensity, and the value of the problem to the original non-Markovian processes and the value of the primal problem. We then formulate continuous-time deterministic policy gradient learning on the Markovianized approximation…
*Auto-collected on 2026-08-21*