论文概要
研究领域: CV 作者: Jiaju Han, Ma Yaqi, Yahui Chai 发布时间: 2025-07-09 arXiv: 2507.06827
English Summary
Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery. Yet, most remote-sensing vision-language resources and models focus on visible-band semantics, leaving infrared vision-language understanding underexplored. The authors introduce MonoIR-RS, a large-scale infrared remote-sensing vision-language dataset and benchmark that couples IR-aware data construction with CLIP-style contrastive adaptation and VLM instruction tuning.
Built from the same source pool and split as FusionRS, MonoIR-RS retains the infrared image as the model-facing modality, yielding 600,000 synthesized infrared images and 59,032 retained IR-aware caption records. The model experiments use this retained language-supervision subset, whose captions are rewritten around grayscale structure and infrared-style contrast rather than RGB appearance.
Key Findings
- Synthesized infrared images in MonoIR-RS are substantially closer to real thermal imagery than simple grayscale conversion, evaluated on the AVIID benchmark.
- The authors fine-tuned five CLIP backbones and six VLM backbones, calibrating them against zero-shot behavior.
- IR-aware adaptation improves CLIP average recall by up to 12.8 points.
- VLM caption infrared-cue coverage is driven to 100%, while residual RGB color leakage drops to near zero.
- arXiv: https://arxiv.org/abs/2507.06827
By isolating the infrared modality from RGB-IR bimodal learning, MonoIR-RS provides a controlled, reproducible testbed for aligning infrared remote-sensing evidence with language.
Links
*Auto-collected on 2026-07-09*