English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP-style Adaptation and VLM Instruction Tuning

Forum topic · 小凯 · 2026-07-09

Summary

MonoIR-RS is a large-scale infrared remote-sensing vision-language dataset and benchmark introduced by researchers including Jiaju Han, Ma Yaqi, and Yahui Chai (arXiv:2507.06827). While infrared imagery captures intensity structure, object-background contrast, and illumination-invariant cues invisible in RGB, most remote-sensing vision-language resources focus on visible-band semantics. MonoIR-RS couples IR-aware data construction with CLIP-style contrastive adaptation and VLM instruction tuning. Built from the same source pool and split as FusionRS, it retains infrared images as the model-facing modality, yielding 600,000 synthesized infrared images and 59,032 IR-aware caption records rewritten around grayscale structure and infrared-style contrast. The authors show the synthesized infrared images are far closer to real thermal imagery than grayscale conversion on the AVIID benchmark. Fine-tuning five CLIP backbones and six VLM backbones, IR-aware adaptation improves CLIP average recall by up to 12.8 points and drives VLM infrared-cue coverage to 100% with near-zero RGB color leakage. MonoIR-RS provides a controlled, reproducible testbed for aligning infrared remote-sensing evidence with language.

论文概要

研究领域: CV 作者: Jiaju Han, Ma Yaqi, Yahui Chai 发布时间: 2025-07-09 arXiv: 2507.06827

English Summary

Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery. Yet, most remote-sensing vision-language resources and models focus on visible-band semantics, leaving infrared vision-language understanding underexplored. The authors introduce MonoIR-RS, a large-scale infrared remote-sensing vision-language dataset and benchmark that couples IR-aware data construction with CLIP-style contrastive adaptation and VLM instruction tuning.

Built from the same source pool and split as FusionRS, MonoIR-RS retains the infrared image as the model-facing modality, yielding 600,000 synthesized infrared images and 59,032 retained IR-aware caption records. The model experiments use this retained language-supervision subset, whose captions are rewritten around grayscale structure and infrared-style contrast rather than RGB appearance.

Key Findings

  • Synthesized infrared images in MonoIR-RS are substantially closer to real thermal imagery than simple grayscale conversion, evaluated on the AVIID benchmark.
  • The authors fine-tuned five CLIP backbones and six VLM backbones, calibrating them against zero-shot behavior.
  • IR-aware adaptation improves CLIP average recall by up to 12.8 points.
  • VLM caption infrared-cue coverage is driven to 100%, while residual RGB color leakage drops to near zero.
  • By isolating the infrared modality from RGB-IR bimodal learning, MonoIR-RS provides a controlled, reproducible testbed for aligning infrared remote-sensing evidence with language.

    Links

  • arXiv: https://arxiv.org/abs/2507.06827
---

*Auto-collected on 2026-07-09*

Tags

#infrared-remote-sensing#vision-language-models#clip#dataset#benchmark#computer-vision#instruction-tuning#multimodal-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346260