English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Depth Estimation

Forum topic · 小凯 · 2026-09-08

Summary

CrossDepth is a computer vision paper (arXiv:2609.05397) by Samer Abualhanud and Max Mehltretter addressing generalizable multi-view depth estimation for autonomous driving. Surround camera rigs offer broad scene coverage, but adjacent images overlap only minimally, forcing most pixel depths to be inferred from monocular appearance cues that may vary across images. The paper targets two key sources of cross-image inconsistency: differing camera intrinsics and the limited receptive field of each image. CrossDepth conditions features on per-pixel camera-aware ray embeddings to handle camera-dependent variations and extends each pixel's receptive field via a geometry-constrained attention mechanism. The work aims to improve robust and consistent 3D environment understanding across multi-view camera systems.

Paper Overview

Field: Computer Vision (CV) Authors: Samer Abualhanud, Max Mehltretter Published: 2026-09-04 arXiv: 2609.05397

Abstract

Reliable 3D understanding of the surrounding environment is a core requirement for autonomous driving. Multi-view surround camera rigs provide broad scene coverage, but the spatially adjacent images typically overlap only minimally. Consequently, the depth of most pixels must be inferred from monocular appearance cues. These cues can appear differently across images and may therefore be interpreted differently by the depth estimation model.

The paper targets two main sources of cross-image inconsistency:

1. Differences in camera intrinsics — addressed by conditioning the features on per-pixel camera-aware ray embeddings, enabling the network to account for camera-dependent variations in monocular cues. 2. Limited receptive field of each image — addressed by extending each pixel's receptive field (via geometry-constrained attention).

Key Points

  • Focus: generalizable multi-view depth estimation for autonomous driving with surround camera rigs.
  • Problem: minimal overlap between adjacent camera views means most depth estimates rely on monocular cues, which can be inconsistent across images.
  • Solution 1: per-pixel camera-aware ray embeddings to compensate for varying camera intrinsics.
  • Solution 2: geometry-constrained attention to extend the receptive field and enforce cross-image consistency.
*Auto-collected on 2026-09-08. The original abstract is truncated in the source post.*

Tags

#paper#arxiv#computer-vision#depth-estimation#multi-view#autonomous-driving#attention-mechanism

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634624