English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Hawkes-CT DDPG: Continuous-Time Reinforcement Learning for Controlled Hawkes Jump-Diffusions in Non-Markovian Settings

Forum topic · 小凯 · 2026-08-21

Summary

Researchers Tomasz R. Bielecki, Thibaut Mastrolia, and Haoze Yan study stochastic control of multivariate Hawkes-driven stochastic differential equations using machine learning in a non-Markovian setting. Because the Hawkes intensity has path-dependent memory, the problem falls outside classical stochastic control theory except for specific Markovian kernels. The authors first develop a finite-dimensional Markovianization procedure that approximates multivariate Hawkes processes with mixtures of exponential kernels, and prove convergence of the approximated process, its intensity, and the problem's value to the original non-Markovian ones. They then build continuous-time deterministic policy gradient learning, called Hawkes-CT DDPG, on the Markovianized approximation. The resulting model-free algorithm solves the non-Markovian Hawkes-driven optimization problem by observing only event times, SDE solution realizations, and a set of decay filters, with Hawkes kernel coefficients remaining unknown. The method is benchmarked against discrete-time reinforcement learning techniques under three kernel types: simple exponential, Erlang, and power-law kernels. Paper: arXiv 2608.19151.

Paper Overview

  • Field: Machine Learning
  • Authors: Tomasz R. Bielecki, Thibaut Mastrolia, Haoze Yan
  • Posted: 2026-08-19
  • arXiv: 2608.19151
  • Introduction

    We study stochastic control of multivariate Hawkes-driven stochastic differential equations with machine learning algorithms in a non-Markovian setting. Due to the path dependence of the memory of the Hawkes intensity, this problem does not fall within classical stochastic control theory outside particular Markovian kernels.

    Key Contributions

    1. Markovianization: A finite-dimensional Markovianization procedure and algorithm that approximate multivariate Hawkes processes with mixtures of exponential kernels. The authors prove the convergence of the Markovianized approximation of the Hawkes process, its intensity, and the value of the problem to the original non-Markovian processes and the value of the primal problem.

    2. Hawkes-CT DDPG: Continuous-time deterministic policy gradient learning formulated on the Markovianized approximation. The proposed model-free algorithm solves the non-Markovian Hawkes-driven optimization problem by observing only:

  • Event times of the process,
  • Realizations of the SDE solution,
  • A selected set of decay filters,
  • while the Hawkes kernel coefficients remain unknown.

    3. Empirical comparison: The continuous-time reinforcement learning approach (Hawkes-CT DDPG) is compared against discrete-time reinforcement learning techniques under three different kernel types:

  • Simple exponential kernel
  • Erlang kernel
  • Power-law kernel
  • Original Abstract (English)

    We study stochastic control of multivariate Hawkes-driven stochastic differential equations with machine learning algorithms in a non-Markovian setting. Due to the path dependence of the memory of the Hawkes intensity, this problem does not fall within classical stochastic control theory outside particular Markovian kernels. We first develop a finite-dimensional Markovianization procedure and algorithm to approximate multivariate Hawkes processes with mixtures of exponential kernels. We prove the convergence of the Markovianized approximation of the Hawkes process, its intensity, and the value of the problem to the original non-Markovian processes and the value of the primal problem. We then formulate continuous-time deterministic policy gradient learning on the Markovianized approximation...

    Links

  • arXiv: <https://arxiv.org/abs/2608.19151>

Tags

#reinforcement-learning#continuous-time-rl#hawkes-process#stochastic-control#machine-learning#non-markovian#policy-gradient#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633742