English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ModMap: Crossmodal Feature Mapping with Cross-View Modulation for Multiview 3D Anomaly Detection

Forum topic · 小凯 · 2026-04-04

Summary

ModMap is a natively multiview and multimodal framework for 3D anomaly detection and segmentation, introduced by Alex Costanzino, Pierluigi Zama Ramirez, and Giuseppe Lisanti in the paper 'Modulate-and-Map' (arXiv 2504.01262, CVPR area). Unlike existing methods that process views independently, ModMap draws on the crossmodal feature mapping paradigm to learn mappings of features across both modalities and viewpoints, explicitly modeling view-dependent relationships through feature-wise modulation. The authors propose a cross-view training strategy that leverages all possible combinations of camera views, enabling effective anomaly scoring via multiview ensembling and aggregation. To handle high-resolution 3D data, they train and publicly release a foundational depth encoder tailored to industrial datasets. Experiments on SiM3D—the first benchmark to introduce a multiview, multimodal setup for 3D anomaly detection and segmentation—show that ModMap outperforms prior methods by a large margin, achieving state-of-the-art performance. The released depth encoder and framework offer practical tools for industrial 3D inspection.

Paper Overview

  • Field: Computer Vision (3D anomaly detection and segmentation)
  • Authors: Alex Costanzino, Pierluigi Zama Ramirez, Giuseppe Lisanti
  • Published: 2025-04-01
  • arXiv: 2504.01262
  • Abstract

    We present ModMap, a natively multiview and multimodal framework for 3D anomaly detection and segmentation. Unlike existing methods that process views independently, our method draws inspiration from the crossmodal feature mapping paradigm to learn to map features across both modalities and views, while explicitly modelling view-dependent relationships through feature-wise modulation.

    We introduce a cross-view training strategy that leverages all possible view combinations, enabling effective anomaly scoring through multiview ensembling and aggregation. To process high-resolution 3D data, we train and publicly release a foundational depth encoder tailored to industrial datasets.

    Experiments on SiM3D, a recent benchmark that introduces the first multiview and multimodal setup for 3D anomaly detection and segmentation, show that ModMap outperforms previous methods by a large margin, achieving state-of-the-art performance.

    Key Contributions

  • A crossmodal feature mapping approach extended across both modalities and camera views.
  • Feature-wise modulation to explicitly model view-dependent relationships.
  • A cross-view training strategy using all possible view combinations, with multiview ensembling and aggregation for anomaly scoring.
  • A publicly released foundational depth encoder designed for high-resolution industrial 3D data.
  • State-of-the-art results on SiM3D, the first multiview multimodal 3D anomaly detection benchmark.
---

*Auto-collected on 2026-04-04.*

Tags

#3d-anomaly-detection#multimodal#computer-vision#depth-estimation#siM3d#arxiv#industrial-inspection

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169522