English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SceneBind: Binding What and Where Across Vision, Audio and Language

Forum topic · 小凯 · 2026-07-18

Summary

SceneBind is an omni-modal scene representation developed by researchers including Mingfei Chen and Eli Shlizerman that jointly captures semantic and 3D spatial understanding across vision, audio, and language (arXiv:2607.15265). While existing omni-modal encoders excel at instance-level semantics—identifying what is present in a scene—they often lack explicit spatial structure indicating where objects are. SceneBind bridges this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots that explicitly capture object-level semantics, spatial attributes, and uncertainty. The authors also introduce SceneBind Matching, a scheme integrating global scene similarity with object alignment to support cross-modal scene retrieval and object grounding. The work includes a newly curated real-world binaural audio-visual dataset with structured semantic and spatial annotations, plus a training protocol for cross-modal alignment. SceneBind is compatible with large-scale pretrained semantic encoders, adding lightweight spatial modeling with few extra tokens, and achieves state-of-the-art scene and spatial retrieval performance with strong zero-shot transfer to tasks like audio-visual grounding.

Paper Overview

  • Field: Computer Vision
  • Authors: Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman
  • Published: 2026-07-16
  • arXiv: 2607.15265
  • Abstract (English, from source)

    We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further propose SceneBind Matching, a semantic-spatial matching scheme that integrates global scene similarity with object alignment, supporting cross-modal scene retrieval and object grounding. To train and evaluate SceneBind, the authors curate a novel real-world binaural audio-visual dataset with structured semantic and spatial annotations and propose a training protocol for cross-modal alignment of semantic and spatial signals.

    Key Contributions

  • Semantic-spatial scene representation: combines a global semantic embedding with object-centric semantic-spatial slots, capturing semantics, spatial attributes, and uncertainty.
  • SceneBind Matching: integrates global scene similarity with object alignment for cross-modal scene retrieval and object grounding.
  • New dataset: a real-world binaural audio-visual dataset with structured semantic and spatial annotations.
  • Efficiency: compatible with large-scale pretrained semantic encoders; adds lightweight spatial modeling with only a few extra tokens.
  • Results: state-of-the-art performance on scene and spatial retrieval, with strong zero-shot transfer to downstream tasks such as audio-visual grounding.
---

*Auto-collected on 2026-07-18*

Tags

#scenebind#multimodal-learning#audio-visual#scene-retrieval#3d-spatial-understanding#object-grounding#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178433592