[论文] Forecasting from Counterfactual Simulator Rollouts: A Sim2Real Evaluat...

研究领域: ML 作者: Angel Wang, Dominique Perrault-Joncas, Alvaro Maggiar, Dean Foster, Carson Eisenach 发布时间: 2026-10-02 arXiv: 2610.03662

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Angel Wang, Dominique Perrault-Joncas, Alvaro Maggiar, Dean Foster, Carson Eisenach 发布时间: 2026-10-02 arXiv: 2610.03662

中文摘要

部署新的决策策略会为预测模型带来冷启动问题——其预测目标依赖于策略的动作:历史观测反映的是先前策略,而新策略下的真实观测尚未获得。仿真通过让目标策略在反事实场景中推演,并用所得轨迹学习系统对这些控制的响应,为解决这一缺口提供了途径。这种仿真训练的模型向现实(Sim2Real)的迁移可以通过过去的真实部署数据来回测评估。利用两个真实世界的库存控制部署,我们从三个角度评估这一过程:仿真器保真度、对真实行为的零样本迁移,以及随真实目标策略观测积累的适应过程。仿真训练的预测器比同一架构在历史真实数据上训练的模型取得更低的点估计平均绝对百分比误差(MAPE),在研究 1 中降低 1.2-3.1 个百分点,在研究 2 中降低 12.5-18.7 个百分点。部署后,使用早期真实观测的轻量校准进一步将误差降低多达 2.5 个百分点。这些结果提供了实证证据:仿真器生成的反事实数据可以支持新策略下的冷启动预测,所得模型可随真实部署数据的积累进一步优化。

原文摘要

Deploying a new decision policy creates a cold-start problem for prediction models whose targets depend on the policy's actions: historical observations reflect earlier policies, while real observations under the new policy are not yet available. Simulation offers a way to address this gap by rolling out the target policy across counterfactual scenarios and using the resulting trajectories to learn how the system responds to those controls. The simulation-to-reality (Sim2Real) transfer of this simulator-trained model can then be backtested by evaluating it against real observations from past deployments. Using two real-world inventory-control deployments, we evaluate this process from three angles: simulator fidelity, zero-shot transfer to real behavior, and adaptation as real target-polic...


*自动采集于 2026-10-06*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens