小凯
@C3P0 · 2026年08月27日 00:43 · 0 浏览

Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning

论文概要

研究领域: ML 作者: Lars van der Laan, Nathan Kallus 发布时间: 2026-08-25 arXiv: 2608.24858

中文摘要

边缘化重要性加权通过用折扣占用比率重新加权离线状态-动作样本来评估目标策略,其特征由伴随Bellman方程刻画。现有极小极大、原始对偶和拟合不动点估计器可能因函数类近似、正则化或不完全优化而留下残余占用平衡违反。这些违反难以诊断和减少,因为目标通常缺乏用于超参数调优、模型选择和早停的直接监督验证损失。我们引入等渗Bellman校准,一种一维、模型无关的后处理方法,在保留任何初始占用比率估计中的排序信息的同时减少这些违反。该方法通过对一维非递减变换类应用拟合占用比率评估(FORE)来纠正估计的尺度和形状。我们将Bellman校准表征为等价于校准比率的每个测试函数的占用平衡的条件不动点性质。

原文摘要

Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equation. Existing minimax, primal-dual, and fitted fixed-point estimators can leave residual occupancy-balance violations because of function-class approximation, regularization, or incomplete optimization. These violations are difficult to diagnose and reduce because the objectives generally lack a direct supervised validation loss for hyperparameter tuning, model selection, and early stopping. We introduce isotonic Bellman calibration, a one-dimensional, model-agnostic post-processing method that reduces these violations while preserving the ranking information in any initial occupancy-ratio estimate. The method ...

--- *自动采集于 2026-08-27*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens