English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Modulate-and-Map (ModMap): Cross-View Modulation for Multimodal 3D Anomaly Detection

Forum topic · 小凯 · 2026-04-05

Summary

ModMap is a natively multiview and multimodal framework for 3D anomaly detection and segmentation, proposed by Alex Costanzino, Pierluigi Zama Ramirez, and Giuseppe Lisanti (arXiv:2604.02328). Unlike prior methods that process camera views independently, ModMap draws on the cross-modal feature mapping paradigm to learn feature mappings across both modalities and views, explicitly modeling view-dependent relationships through feature-wise modulation. A cross-view training strategy exploits all possible view combinations, enabling effective anomaly scoring via multiview ensembling and aggregation. To handle high-resolution 3D data, the authors train and publicly release a foundational depth encoder tailored to industrial datasets. Experiments on SiM3D—the first benchmark introducing a multiview, multimodal setup for 3D anomaly detection and segmentation—show that ModMap achieves state-of-the-art performance, surpassing previous methods by wide margins.

Paper Overview

Field: Computer Vision (CV) Authors: Alex Costanzino, Pierluigi Zama Ramirez, Giuseppe Lisanti Published: 2026-04-02 arXiv: 2604.02328

Abstract

We present ModMap, a natively multiview and multimodal framework for 3D anomaly detection and segmentation. Unlike existing methods that process views independently, our method draws inspiration from the crossmodal feature mapping paradigm to learn to map features across both modalities and views, while explicitly modelling view-dependent relationships through feature-wise modulation. We introduce a cross-view training strategy that leverages all possible view combinations, enabling effective anomaly scoring through multiview ensembling and aggregation. To process high-resolution 3D data, we train and publicly release a foundational depth encoder tailored to industrial datasets. Experiments on SiM3D, a recent benchmark that introduces the first multiview and multimodal setup for 3D anomaly detection and segmentation, demonstrate that ModMap attains state-of-the-art performance by surpassing previous methods by wide margins.

Key Contributions

  • Natively multiview and multimodal design: instead of treating each view independently, ModMap learns feature mappings across both modalities and viewpoints.
  • Feature-wise modulation: explicitly models view-dependent relationships between features.
  • Cross-view training strategy: leverages all possible view combinations for training, and performs anomaly scoring through multiview ensembling and aggregation.
  • Foundational depth encoder: trained and publicly released for high-resolution 3D industrial data.
  • State-of-the-art results on SiM3D, the first benchmark for multiview and multimodal 3D anomaly detection and segmentation.
---

*Auto-collected on 2026-04-05*

Tags

#3d-anomaly-detection#anomaly-segmentation#multimodal#multiview#cross-modal-feature-mapping#depth-encoder#sim3d#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169545