Loading...
正在加载...
请稍候

Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning

小凯 (C3P0) 2026年08月27日 00:43

论文概要

研究领域: ML
作者: Lars van der Laan, Nathan Kallus
发布时间: 2026-08-25
arXiv: 2608.24858

中文摘要

边缘化重要性加权通过用折扣占用比率重新加权离线状态-动作样本来评估目标策略,其特征由伴随Bellman方程刻画。现有极小极大、原始对偶和拟合不动点估计器可能因函数类近似、正则化或不完全优化而留下残余占用平衡违反。这些违反难以诊断和减少,因为目标通常缺乏用于超参数调优、模型选择和早停的直接监督验证损失。我们引入等渗Bellman校准,一种一维、模型无关的后处理方法,在保留任何初始占用比率估计中的排序信息的同时减少这些违反。该方法通过对一维非递减变换类应用拟合占用比率评估(FORE)来纠正估计的尺度和形状。我们将Bellman校准表征为等价于校准比率的每个测试函数的占用平衡的条件不动点性质。

原文摘要

Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equation. Existing minimax, primal-dual, and fitted fixed-point estimators can leave residual occupancy-balance violations because of function-class approximation, regularization, or incomplete optimization. These violations are difficult to diagnose and reduce because the objectives generally lack a direct supervised validation loss for hyperparameter tuning, model selection, and early stopping. We introduce isotonic Bellman calibration, a one-dimensional, model-agnostic post-processing method that reduces these violations while preserving the ranking information in any initial occupancy-ratio estimate. The method ...


自动采集于 2026-08-27

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录