[论文] Trust Guided Decision Transformer
研究领域: ML 作者: Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath 发布时间: 2026-09-25 arXiv…
论文概要
研究领域: ML 作者: Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath 发布时间: 2026-09-25 arXiv: 2609.31586
中文摘要
决策 Transformer 在长程 rollout 中性能会退化,因为条件上下文逐渐漂移出训练分布。我们证明,这种漂移可以通过模型自身的下一步状态预测误差直接观测到——该误差在 rollout 过程中上升并保持高位,从而给出上下文何时变得不可靠的直接信号。我们提出信任引导决策 Transformer(TGDT),它在施加价值引导之前先选择上下文。每一步,TGDT 用滚动的下一步状态预测误差评估最近若干上下文后缀,并通过 split conformal prediction 在留出的离线数据上校准阈值,只保留误差在校准阈值内的后缀,然后用冻结的 critic 在受信任的后缀中选择价值最高的动作。这颠倒了仅用价值进行弹性选择的次序——后者可能选出一个由模型自己已标记为不可靠的上下文所生成的动作。在 D4RL 导航与运动任务上的实验表明,状态预测、critic 引导与硬上下文重置各只解决了问题的一部分。TGDT 减少了持续性高误差运行,回报超过原始决策 Transformer、基于重置的上下文控制以及仅按价值选择上下文的方法。
原文摘要
Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the ord...
*自动采集于 2026-09-29*
#论文 #arXiv #ML #小凯