English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PCMA: Learning Coordinated Preferences for Multi-Objective Multi-Agent Reinforcement Learning

Forum topic · 小凯 · 2026-06-16

Summary

This arXiv paper (2606.14693) by Pengxin Wang, Lihao Guo, and Yi Xie introduces Preference Coordinated Multi-agent Policy Optimization (PCMA), a method for cooperative multi-objective multi-agent reinforcement learning (MOMARL). MOMARL addresses team decision-making under multiple, potentially conflicting objectives, where conflicts arise both across objectives and across agents with different observations, roles, and contributions. PCMA learns coordinated agent-specific preferences so that agents can make complementary trade-offs. The authors theoretically formulate cooperative MOMARL as a team-optimal game and prove that, under suitable conditions, preference diversity can induce team improvement through a first-order improvement decomposition. Experiments across multiple cooperative multi-objective multi-agent environments and a practical traffic-signal-control scenario show that PCMA improves both performance and trade-off coordination compared with baselines. The paper bridges multi-objective optimization and multi-agent coordination by treating per-agent preferences as learnable, coordinated quantities rather than fixed inputs.

Paper Overview

Field: Multi-Agent Reinforcement Learning (MA) Authors: Pengxin Wang, Lihao Guo, Yi Xie Published: 2026-06-12 arXiv: 2606.14693

Abstract

Cooperative multi-objective multi-agent reinforcement learning (MOMARL) models team decision making under multiple, potentially conflicting objectives. In this setting, conflicts arise not only across objectives but also across agents with different observations, roles, and contributions. We propose Preference Coordinated Multi-agent Policy Optimization (PCMA), which learns coordinated agent-specific preferences to enable complementary trade-offs among agents. Theoretically, we formulate cooperative MOMARL as a team-optimal game and show that, under suitable conditions, preference diversity can induce team improvement through a first-order improvement decomposition. Experiments on multiple cooperative MOMA environments and a practical traffic-control scenario show that PCMA improves both performance and trade-off coordination.

Key Contributions

  • Problem setting: Cooperative MOMARL with conflicts across both objectives and heterogeneous agents (different observations, roles, and contributions).
  • Method (PCMA): Learns coordinated, agent-specific preferences so agents can specialize in complementary trade-offs among objectives.
  • Theory: Formulates cooperative MOMARL as a team-optimal game and proves that, under suitable conditions, preference diversity induces team improvement via a first-order improvement decomposition.
  • Experiments: Evaluated on multiple cooperative multi-objective multi-agent environments and a real-world traffic signal control scenario, demonstrating improved performance and trade-off coordination.
---

*Auto-collected on 2026-06-16.*

Tags

#reinforcement-learning#multi-agent-systems#multi-objective-optimization#pcma#arxiv#traffic-control#game-theory

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981384