Paper Overview
Research area: Computer Vision Authors: Jiaju Han, Ma Yaqi, Yahui Chai Release date: 2025-07-09 arXiv: 2507.06827
Abstract
Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery. Yet, most remote-sensing vision-language resources and models focus on visible-band semantics, leaving infrared vision-language understanding underexplored.
MonoIR-RS is a large-scale infrared remote-sensing vision-language dataset and benchmark that couples IR-aware data construction with CLIP-style contrastive adaptation and VLM instruction tuning. Built from the same source pool and split as FusionRS, MonoIR-RS retains the infrared image as the model-facing modality, yielding 600,000 synthesized infrared images and 59,032 retained IR-aware caption records. The model experiments use this retained language-supervision subset, whose captions are rewritten to supervise grayscale structure and infrared-style contrast rather than RGB appearance.
Key Findings
- Synthesized infrared images on the AVIID benchmark are markedly closer to real thermal imagery than simple grayscale conversion.
- Five CLIP backbones and six VLM backbones were fine-tuned and calibrated against zero-shot behavior.
- IR-aware adaptation improves CLIP average recall by up to 12.8 points.
- VLM caption IR-cue coverage is driven to 100%, while residual RGB color leakage is reduced to near zero.
- By isolating the infrared modality from RGB-IR bimodal learning, MonoIR-RS provides a controlled, reproducible testbed for aligning infrared remote-sensing evidence with language.