English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SceneBind: Binding What and Where Across Vision, Audio and Language

Forum topic · 小凯 · 2026-07-18

Summary

SceneBind (arXiv:2607.15265) is an omni-modal scene representation that jointly provides semantic and 3D spatial understanding across vision, audio, and language. While existing omni-modal encoders capture instance-level semantics (what is present), they often lack explicit spatial structure (where it is). SceneBind fills this gap by representing each scene as a semantic-spatial entity: a global semantic embedding combined with object-centric semantic-spatial slots that explicitly capture object-level semantics, spatial attributes, and uncertainty. The authors also introduce SceneBind Matching, a scheme integrating global scene similarity with object alignment to support cross-modal scene retrieval and object grounding. For training and evaluation, the team curated a novel real-world binaural audio-visual dataset with structured semantic and spatial annotations, plus a training protocol aligning semantic and spatial signals across modalities. SceneBind is compatible with large pretrained semantic encoders, adding lightweight spatial modeling with only a few extra tokens. It achieves state-of-the-art performance in scene and spatial retrieval and shows strong zero-shot transfer to downstream tasks such as audio-visual grounding.

Paper Overview

Research Area: CV Authors: Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman Published: 2026-07-16 arXiv: 2607.15265

Abstract

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty.

We further propose SceneBind Matching, a semantic-spatial matching scheme that integrates global scene similarity with object alignment, supporting cross-modal scene retrieval and object grounding.

To train and evaluate SceneBind, the authors curated a novel real-world binaural audio-visual dataset with structured semantic and spatial annotations, and proposed a training protocol that aligns semantic and spatial signals across modalities.

Key Highlights

  • Joint semantic + spatial representation: combines a global scene embedding with object-centric semantic-spatial slots capturing object semantics, spatial attributes, and uncertainty.
  • SceneBind Matching: integrates global scene similarity with object alignment for cross-modal scene retrieval and object grounding.
  • New binaural audio-visual dataset: real-world data with structured semantic and spatial annotations.
  • Lightweight: compatible with large-scale pretrained semantic encoders, adding only a few extra tokens for spatial modeling.
  • Strong results: state-of-the-art on scene and spatial retrieval, with robust zero-shot transfer to downstream tasks such as audio-visual localization/grounding.
---

*Auto-collected on 2026-07-18*

Tags

#scenebind#multimodal-representation#audio-visual#3d-spatial-understanding#cross-modal-retrieval#object-grounding#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178433583