Paper Overview
Research Area: Computer Vision (CV) Authors: Hao Dong, Hongzhao Li, Shupan Li, Muhammad Haris Khan et al. arXiv: 2605.06643
Abstract
Although Multimodal Domain Generalization (MMDG) is increasingly popular for improving model robustness, it remains unclear whether reported performance gains reflect genuine algorithmic progress or artifacts of inconsistent evaluation protocols. Existing research is highly fragmented, with studies differing significantly in datasets, modality configurations, and experimental settings. Moreover, current benchmarks focus mainly on action recognition and often overlook critical real-world challenges such as input corruption, missing modalities, and model trustworthiness. This lack of standardization obscures reliable assessment of progress in the field.
MMDG-Bench
To address this, the authors introduce MMDG-Bench, the first unified and comprehensive MMDG benchmark:
- Standardized evaluation on six datasets spanning three distinct tasks: action recognition, mechanical fault diagnosis, and sentiment analysis
- Six modality combinations and nine representative methods under multiple evaluation settings
- Beyond standard accuracy, systematic evaluation of corruption robustness, missing-modality generalization, misclassification detection, and out-of-distribution detection
- 7,402 neural networks trained across 95 unique cross-domain tasks
Key Findings
1. Under fair comparison, recent specialized MMDG methods offer only marginal improvements over the ERM baseline. 2. No single method consistently outperforms others across datasets or modality combinations. 3. A significant gap to upper-bound performance remains, indicating MMDG is far from solved. 4. Three-modality fusion does not consistently outperform the strongest two-modality configurations. 5. All evaluated methods show significant degradation under corruption and missing-modality scenarios, and some methods further compromise model trustworthiness.
---
*Source: arXiv:2605.06643, auto-collected on 2026-05-10.*