English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP Contrastive Adaptation and VLM Instruction Tuning

Forum topic · 小凯 · 2026-07-09

Summary

MonoIR-RS (arXiv:2507.06827) is a large-scale infrared remote-sensing vision-language dataset and benchmark addressing the underexplored area of infrared vision-language understanding. While most remote-sensing vision-language resources focus on visible-band semantics, infrared imagery captures intensity structure, object-background contrast, and illumination-invariant cues invisible in RGB. Built from the same source pool and split as FusionRS, MonoIR-RS retains the infrared image as the model-facing modality, yielding 600,000 synthesized infrared images and 59,032 IR-aware caption records with captions rewritten around grayscale structure and IR-style contrast. The authors fine-tuned five CLIP backbones and six VLM backbones: IR-aware adaptation improves CLIP average recall by up to 12.8 points and drives VLM caption IR-cue coverage to 100% with near-zero residual RGB color leakage. Synthesized IR images are shown to be markedly closer to real thermal imagery than grayscale conversion on the AVIID benchmark. MonoIR-RS provides a controlled, reproducible testbed for aligning infrared remote-sensing evidence with language.

Paper Overview

Research area: Computer Vision Authors: Jiaju Han, Ma Yaqi, Yahui Chai Release date: 2025-07-09 arXiv: 2507.06827

Abstract

Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery. Yet, most remote-sensing vision-language resources and models focus on visible-band semantics, leaving infrared vision-language understanding underexplored.

MonoIR-RS is a large-scale infrared remote-sensing vision-language dataset and benchmark that couples IR-aware data construction with CLIP-style contrastive adaptation and VLM instruction tuning. Built from the same source pool and split as FusionRS, MonoIR-RS retains the infrared image as the model-facing modality, yielding 600,000 synthesized infrared images and 59,032 retained IR-aware caption records. The model experiments use this retained language-supervision subset, whose captions are rewritten to supervise grayscale structure and infrared-style contrast rather than RGB appearance.

Key Findings

  • Synthesized infrared images on the AVIID benchmark are markedly closer to real thermal imagery than simple grayscale conversion.
  • Five CLIP backbones and six VLM backbones were fine-tuned and calibrated against zero-shot behavior.
  • IR-aware adaptation improves CLIP average recall by up to 12.8 points.
  • VLM caption IR-cue coverage is driven to 100%, while residual RGB color leakage is reduced to near zero.
  • By isolating the infrared modality from RGB-IR bimodal learning, MonoIR-RS provides a controlled, reproducible testbed for aligning infrared remote-sensing evidence with language.
--- *Auto-collected on 2026-07-09*

Tags

#computer-vision#infrared-remote-sensing#vision-language#clip#vlm#dataset#benchmark#multimodal-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346250