Overview
- Field: Machine Learning
- Author: Domagoj Herceg
- Published: 2026-06-26
- arXiv: 2606.28281
- Applies PAC-Bayesian theory to learning-based control with quadratic trajectory costs (unbounded, non-Lipschitz losses).
- Uses System Level Synthesis (SLS) parameterization to expose the closed-loop trajectory map and enable explicit certification.
- Provides PAC-Bayes-Chernoff certificates for posteriors over feasible closed-loop responses.
- Derives an exact one-sided Gaussian transform and tractable quadratic upper bounds for Gaussian disturbances with arbitrary covariance.
- Shows certificates transfer from the posterior to the deterministic mean response; includes a data-driven deployment bound.
- Experiments on a double integrator demonstrate improved held-out cost and reduced closed-loop sensitivity in low-data settings.
English Abstract
PAC-Bayesian bounds provide finite-sample guarantees for data-dependent randomized predictors, but applying them to learning-based control is difficult because the natural objective is a quadratic trajectory cost. Such losses are unbounded, non-Lipschitz, and lead to response-dependent Chernoff terms. We employ System Level Synthesis parameterization, which exposes the closed-loop trajectory map of a linear system directly and makes the quadratic control loss amenable to explicit certification. Moreover, we provide a set of PAC-Bayes-Chernoff certificates for posterior distributions over feasible closed-loop responses. For Gaussian disturbance trajectories with arbitrary covariance, we derive an exact one-sided Gaussian transform and a tractable quadratic upper bound expressed through closed-loop sensitivity quantities. Although PAC-Bayes certifies non-degenerate posteriors, the convex quadratic form of the SLS loss transfers certificates to the posterior mean response. We propose a deterministic mean-response deployment result particularly suited for control while retaining the randomized posterior in the bound, and provide a data-driven bound for this deployment. Numerical experiments on a double integrator show that the algorithm acts as a sensitivity-aware finite-sample regularizer, improving held-out costs and reducing closed-loop sensitivity in low-data regimes.