English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

Forum topic · 小凯 · 2026-08-14

Summary

This post introduces an arXiv paper (2508.03417) by Ebenezer Gelo, Geraud Nangue Tasse, and Steven James on safe offline reinforcement learning. Standard approaches assume dense per-step cost annotations, but in practice supervisors only provide trajectory-level stopping feedback: a binary signal at the first unsafe transition, with no per-step attribution. The authors frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse stopping feedback into dense per-step costs via return decomposition, then trains a constrained offline policy on the augmented dataset. They prove that return-equivalent redistributions preserve the feasible policy set and optimal Lagrangian in constrained MDPs, establishing that the transformation is theoretically lossless while yielding better-conditioned cost critic learning in practice. Experiments on highway driving and robotic manipulation show significantly reduced violation rates compared to sparse and classifier-based baselines, with robustness to heterogeneous dataset composition and label noise.

Paper Overview

Field: Machine Learning Authors: Ebenezer Gelo, Geraud Nangue Tasse, Steven James Published: 2026-08-13 arXiv: 2508.03417

Summary

Safe offline reinforcement learning typically assumes access to dense per-step cost annotations, but in practice supervisors only provide trajectory-level stopping feedback: a binary signal at the first unsafe transition, with no per-step attribution. The authors frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse stopping feedback into dense per-step costs through return decomposition, then trains a constrained offline policy on the augmented dataset.

The authors prove that return-equivalent redistributions preserve the feasible policy set and the optimal Lagrangian in constrained MDPs (CMDPs), establishing that the transformation is theoretically lossless, while in practice producing better-conditioned cost critic learning.

Experiments on highway driving and robotic manipulation demonstrate significantly lower violation rates compared to sparse and classifier-based baselines, along with robustness to heterogeneous dataset composition and label noise.

Original Abstract

Safe offline RL typically assumes access to dense per-step cost annotations... (see full paper at arXiv:2508.03417)

--- *Auto-collected on 2026-08-14*

Tags

#reinforcement-learning#safe-rl#offline-rl#cost-inference#cmdp#credit-assignment#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633455