English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Zooming: GeoMTVR Dataset and GeoLens Model for Multi-Tool Visual Reasoning in Ultra-High-Resolution Remote Sensing

Forum topic · 小凯 · 2026-07-30

Summary

This paper addresses the limitations of multimodal large language models (MLLMs) on ultra-high-resolution (UHR) remote-sensing imagery, where task-relevant evidence is sparse, local, and spatially dispersed across massive visual contexts. A pilot study on XLRS-Bench shows that zoom-in tools only partially help: they resolve easy and medium tasks with locally recoverable evidence but saturate on hard cases requiring global search, multi-region comparison, path planning, or dispersed-evidence reasoning. The authors introduce GeoMTVR, a large-scale geospatial multi-tool visual reasoning dataset built from wide-area satellite imagery, containing 13K UHR VQA samples with interleaved reasoning traces, diverse visual tool calls, and returned visual observations. This enables models to learn question decomposition, tool selection, region inspection, object-level localization, auxiliary visual reasoning, and cross-tool evidence integration. Beyond supervised fine-tuning, they propose a tool-attention-focused reinforcement learning algorithm that concentrates optimization on key decisions: when to call a tool, which tool to choose, where to apply it, and how to interpret its output. Combining SFT and RL on GeoMTVR yields GeoLens, a multi-tool visual reasoning MLLM that consistently outperforms direct inference and single-tool zooming baselines in accuracy, evidence localization, and tool-use trajectory quality.

Overview

Field: Computer Vision (CV) Authors: Fengxiang Wang, Jiangnan Huang, Mingshuo Chen, Yueying Li, Yang Shi, Junwei Luo, Haoyu Wang, Yansheng Li, Jing Zhang, Haiyan Zhao, Wenjing Yang arXiv: 2607.25993

Motivation

Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal large language models (MLLMs): task-relevant evidence is often sparse, local, and spatially dispersed across extremely large visual contexts.

A natural solution is to equip MLLMs with zoom-in tools for active local inspection. However, a pilot study on XLRS-Bench reveals that zoom-in is only partially effective:

  • It resolves easy and medium-level tasks with locally recoverable evidence.
  • It saturates on hard cases requiring global search, multi-region comparison, path planning, or dispersed-evidence reasoning.
  • Contributions

    1. GeoMTVR Dataset

    Moving beyond single-tool zooming, the authors introduce GeoMTVR, a large-scale geospatial multi-tool visual reasoning dataset built from wide-area satellite imagery. It contains 13K UHR VQA samples featuring:

  • Interleaved reasoning trajectories
  • Diverse visual tool calls with returned visual observations
  • Training signals for question decomposition, tool selection, region inspection, object-level localization, auxiliary visual reasoning, and cross-tool evidence integration
  • 2. Tool-Attention-Focused Reinforcement Learning

    Beyond supervised fine-tuning (SFT), the authors propose an RL algorithm that concentrates optimization on key tool-use decisions:

  • When to invoke a tool
  • Which tool to select
  • Where to apply it
  • How to interpret tool outputs
  • 3. GeoLens Model

    By combining SFT on GeoMTVR with the proposed RL algorithm, the authors develop GeoLens, a multi-tool visual reasoning MLLM for UHR remote sensing.

    Results

    Experiments show that GeoLens consistently outperforms both direct inference and single-tool zooming baselines, achieving:

  • Stronger task accuracy
  • Better evidence localization
  • More effective tool-use trajectories

Tags

#computer-vision#remote-sensing#multimodal-llm#visual-reasoning#reinforcement-learning#tool-use#dataset#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503804