English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Bi-CMPStereo: Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo Matching

Forum topic · 小凯 · 2026-04-18

Summary

Researchers Ninghui Xu, Fabio Tosi, and Lihui Wang present Bi-CMPStereo, a bidirectional cross-modal prompting framework for event-frame asymmetric stereo matching, published as arXiv:2504.13101 (April 2025). The work addresses a key limitation in 3D perception: conventional frame cameras capture rich context but suffer from motion blur and limited temporal resolution in dynamic scenes, while event cameras offer high dynamic range but sparse data. Although combining the two modalities enables reliable depth perception under fast motion and challenging illumination, the modality gap tends to marginalize domain-specific cues essential for cross-modal stereo matching. Bi-CMPStereo exploits semantic and structural features from both domains, learning finely aligned stereo representations in a target canonical space and integrating complementary information by projecting each modality into both event and frame domains. Experiments show the method significantly outperforms state-of-the-art approaches in accuracy and generalization.

Bi-CMPStereo: Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo

Research area: Computer Vision Authors: Ninghui Xu, Fabio Tosi, Lihui Wang Published: 2025-04-17 arXiv: 2504.13101

Overview

Conventional frame-based cameras capture rich contextual information but suffer from limited temporal resolution and motion blur in dynamic scenes. Event cameras offer an alternative visual representation with higher dynamic range free from such limitations. The complementary characteristics of the two modalities make event-frame asymmetric stereo promising for reliable 3D perception under fast motion and challenging illumination. However, the modality gap often leads to marginalization of domain-specific cues essential for cross-modal stereo matching.

In this paper, the authors introduce Bi-CMPStereo, a novel bidirectional cross-modal prompting framework that fully exploits semantic and structural features from both domains for robust matching. The approach learns finely aligned stereo representations in a target canonical space and integrates complementary representations by projecting each modality into both the event domain and the frame domain.

Key Contributions

  • A bidirectional cross-modal prompting framework that leverages semantic and structural cues from both event and frame modalities.
  • Finely aligned stereo representation learning in a canonical target space.
  • Integration of complementary information via projecting each modality into both domains.
  • Results

    Extensive experiments demonstrate that Bi-CMPStereo significantly outperforms existing state-of-the-art methods in both accuracy and generalization for event-frame stereo matching.

    Links

  • Paper: https://arxiv.org/abs/2504.13101
--- *Auto-collected on 2026-04-18.*

Tags

#computer-vision#event-camera#stereo-matching#3d-perception#cross-modal#depth-estimation#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618535