[论文] [论文] A Decentralized Partially Observable Team Decision Methodolo...
论文概要 研究领域: RL 作者: Xiaoxing Ren, Thomas Parisini, Andreas A. Malikopoulos 发布时间: 2026-09-22 arXiv: 2609.26783
论文概要
研究领域: RL 作者: Xiaoxing Ren, Thomas Parisini, Andreas A. Malikopoulos 发布时间: 2026-09-22 arXiv: 2609.26783中文摘要
我们研究具有低秩潜在动态与未知系统模型的去中心化部分可观测团队决策问题。所提框架结合团队论等价与低秩模型表示,处理部分可观测 MDP 中的协作决策,且无需先验转移模型。每个成员基于本地私有信息与团队共享的延迟公共信息决策;仅用这些可得信息,各成员学习近似低秩 MDP 并应用最小二乘值迭代计算策略,得到完全去中心化的学习与规划算法——既无中心化协调器,也无中心化训练。我们证明成员侧解逼近中心化团队解:尽管存在部分可观测性、未知动态与延迟公共信息,每个成员仍能恢复近似团队最优策略的相应分量。进一步建立有限样本性能保证,并推导算法的样本复杂度界。原文摘要
We study decentralized partially observable team decision problems with low-rank latent dynamics and unknown system models. The proposed framework combines team-theoretic equivalence with low-rank model representations to address cooperative decision-making in partially observable Markov decision processes without prior knowledge of the transition model. Each team member makes decisions based on local private information and delayed common information shared across the team. Using only this available information, each member learns an approximate low-rank Markov decision process and applies least-squares value iteration to compute its policy. This yields a fully decentralized learning and planning algorithm that requires neither a centralized coordinator nor centralized training. We show that the resulting member-side solutions approximate the centralized team solution: despite partial observability, unknown dynamics, and delayed common information, each member recovers the corresponding component of an approximate team-optimal policy. We further establish finite-sample performance guarantees and derive a corresponding sample-complexity bound for the proposed algorithm.*自动采集于 2026-09-24*
#论文 #arXiv #RL #小凯