English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DAR-Net: Semantically-Aware Diver Activity Recognition for Underwater Human-Robot Collaboration

Forum topic · 小凯 · 2026-06-12

Summary

Researchers Sadman Sakib Enan and Junaed Sattar introduce DAR-Net, a novel transformer-based computer vision framework for recognizing diver activities in underwater scenes, enabling autonomous underwater vehicles (AUVs) to act as effective teammates in high-risk underwater operations. DAR-Net couples transformer-based temporal reasoning with pixel-level scene supervision through a semantically guided, multi-loss training strategy that aligns global activity recognition with local human-robot interaction semantics—a critical factor in low-visibility underwater conditions. To address severe data scarcity in this domain, the authors present the first Underwater Diver Activity (UDA) dataset, a foundational resource containing over 2,600 images annotated with pixel-level masks. Rigorous experiments in controlled environments show that DAR-Net achieves promising accuracy in classifying six distinct diver activities, outperforming state-of-the-art models. The work is available on arXiv (2606.12374) and lays groundwork for smarter, collaborative underwater robotic systems.

Overview

  • Field: Computer Vision
  • Authors: Sadman Sakib Enan, Junaed Sattar
  • Published: 2026-06-10
  • arXiv: 2606.12374

Abstract

Effective multi-human-robot collaboration is essential for expanding human-led operations in the challenging and high-risk underwater environment. For autonomous underwater vehicles (AUVs) to become true teammates, they must be able to comprehend their surroundings and recognize a diver's activities to offer assistance and ensure safety.

Towards this goal, the authors introduce DAR-Net, a novel transformer-based framework that analyzes complex underwater scenes to classify diver activities. The key contribution lies in a semantically guided learning formulation that couples transformer-based temporal reasoning with pixel-level scene supervision. This multi-loss training strategy explicitly aligns global activity recognition with local human-robot interaction semantics, which is particularly critical in low-visibility underwater conditions.

To address the significant challenge of data scarcity in this field, the paper presents the Underwater Diver Activity (UDA) dataset, a foundational resource containing over 2,600 images with pixel-level mask annotations. Through rigorous experimental evaluation in controlled environments, the authors demonstrate that DAR-Net achieves promising accuracy in recognizing six distinct diver activities, outperforming state-of-the-art models. While the dataset provides a crucial baseline, this work serves as a pioneering step, laying the foundation for future research toward smarter, collaborative underwater robotic systems.

Key Contributions

1. DAR-Net: A transformer-based framework for diver activity classification in complex underwater scenes. 2. Semantically guided multi-loss training: Aligns global activity recognition with pixel-level human-robot interaction semantics. 3. UDA dataset: The first underwater diver activity dataset with 2,600+ pixel-level mask annotated images.

*Auto-collected on 2026-06-12.*

Tags

#computer-vision#transformer#underwater-robotics#human-robot-collaboration#activity-recognition#dataset#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981129