English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Decentralized Partially Observable Team Decision-Making with Low-Rank Latent Dynamics

Forum topic · 小凯 · 2026-09-24

Summary

This paper by Xiaoxing Ren, Thomas Parisini, and Andreas A. Malikopoulos (arXiv:2609.26783) studies decentralized team decision-making in partially observable Markov decision processes with low-rank latent dynamics and unknown system models. The framework combines team-theoretic equivalence with low-rank model representations, removing the need for prior knowledge of the transition model. Each team member acts on local private information plus delayed common information shared across the team, learns an approximate low-rank MDP, and applies least-squares value iteration to compute its policy. The result is a fully decentralized learning and planning algorithm requiring neither a centralized coordinator nor centralized training. The authors prove that member-side solutions approximate the centralized team solution, showing that despite partial observability, unknown dynamics, and delayed common information, each member recovers the corresponding component of an approximate team-optimal policy. Finite-sample performance guarantees and a sample-complexity bound are also derived for the proposed algorithm.

We study decentralized partially observable team decision problems with low-rank latent dynamics and unknown system models. The proposed framework combines team-theoretic equivalence with low-rank model representations to address cooperative decision-making in partially observable Markov decision processes without prior knowledge of the transition model.

Each team member makes decisions based on local private information and delayed common information shared across the team. Using only this available information, each member learns an approximate low-rank Markov decision process and applies least-squares value iteration to compute its policy. This yields a fully decentralized learning and planning algorithm that requires neither a centralized coordinator nor centralized training.

We show that the resulting member-side solutions approximate the centralized team solution: despite partial observability, unknown dynamics, and delayed common information, each member recovers the corresponding component of an approximate team-optimal policy. We further establish finite-sample performance guarantees and derive a corresponding sample-complexity bound for the proposed algorithm.

Paper info: arXiv 2609.26783 — Xiaoxing Ren, Thomas Parisini, Andreas A. Malikopoulos

Tags

#reinforcement-learning#decentralized-control#partially-observable-mdp#team-decision-theory#low-rank-dynamics#sample-complexity#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635141