Overview
- Field: Computer Vision
- Authors: Sadman Sakib Enan, Junaed Sattar
- Published: 2026-06-10
- arXiv: 2606.12374
Abstract
Effective multi-human-robot collaboration is essential for expanding human-led operations in the challenging and high-risk underwater environment. For autonomous underwater vehicles (AUVs) to become true teammates, they must be able to comprehend their surroundings and recognize a diver's activities to offer assistance and ensure safety.
Towards this goal, the authors introduce DAR-Net, a novel transformer-based framework that analyzes complex underwater scenes to classify diver activities. The key contribution lies in a semantically guided learning formulation that couples transformer-based temporal reasoning with pixel-level scene supervision. This multi-loss training strategy explicitly aligns global activity recognition with local human-robot interaction semantics, which is particularly critical in low-visibility underwater conditions.
To address the significant challenge of data scarcity in this field, the paper presents the Underwater Diver Activity (UDA) dataset, a foundational resource containing over 2,600 images with pixel-level mask annotations. Through rigorous experimental evaluation in controlled environments, the authors demonstrate that DAR-Net achieves promising accuracy in recognizing six distinct diver activities, outperforming state-of-the-art models. While the dataset provides a crucial baseline, this work serves as a pioneering step, laying the foundation for future research toward smarter, collaborative underwater robotic systems.
Key Contributions
1. DAR-Net: A transformer-based framework for diver activity classification in complex underwater scenes. 2. Semantically guided multi-loss training: Aligns global activity recognition with pixel-level human-robot interaction semantics. 3. UDA dataset: The first underwater diver activity dataset with 2,600+ pixel-level mask annotated images.
*Auto-collected on 2026-06-12.*