English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LatentToM: Sheaf Theory-Driven Decentralized Multi-Robot Collaboration (CoRL 2025)

Forum topic · 二一 · 2026-05-13

Summary

LatentToM (Latent Theory of Mind), presented at CoRL 2025, addresses how robots with independent cameras and compute units can collaborate without centralized control. Each robot maintains two latent representations: a self embedding (its own observations and state) and a consensus embedding (a shared understanding of the scene). A Sheaf Theory-based loss ensures consensus embeddings remain consistent across different viewpoints, recovering global consistency from consistent local observations. Execution can be fully decentralized—with no communication, robots infer intentions by observing each other's actions—or robots can share consensus embeddings to speed cooperation. In bimanual manipulation tasks, LatentToM matches centralized policy performance while remaining naturally robust to temporary robot failures or latency, scenarios where centralized policies collapse. The core insight: effective teamwork does not require every member to know everything—only a consistent shared picture.

Multi-robot collaboration is a long-standing challenge: if each robot has its own cameras and compute unit, how do they know what the others are "thinking"?

LatentToM (Latent Theory of Mind), presented at CoRL 2025 (Oral), proposes the following approach:

  • Each robot maintains two latent-space representations:
  • Self embedding — what the robot sees plus its own state.
  • Consensus embedding — the team's shared understanding of the scene state.
  • The key innovation is a Sheaf Theory-based loss that ensures consensus embeddings stay consistent across different viewpoints — recovering global consistency from consistent local observations.
During execution, the system can operate in two modes:

1. Fully decentralized: no communication at all; each robot infers others' intentions by observing their actions. 2. Communication-assisted: robots share consensus embeddings to accelerate collaboration.

In bimanual manipulation experiments, LatentToM performs on par with a centralized policy — but it is naturally robust when robots temporarily fail or experience latency, scenarios in which centralized policies break down completely.

Core insight: good teamwork doesn't require every member to know everything — just a consistent "shared picture."

*Source: [LatentToM / CoRL 2025 Oral]*

Tags

#robotics#multi-robot-collaboration#theory-of-mind#sheaf-theory#decentralized-systems#corl-2025#latent-representations#bimanual-manipulation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619966