Paper Overview
Field: Computer Vision Authors: Pengcheng Zhou, Xuanyu Liu, Yanchen Yin Published: 2025-07-16 arXiv: 2507.12513
Summary
Recent advances in Vision-Language Models (VLMs) have significantly improved image geo-localization, yet existing models remain susceptible to landmark bias, causing them to overlook geographical cues or form spurious correlations, ultimately resulting in inaccurate localization.
To systematically investigate this issue, the authors design two quantitative metrics, Bias Intensity (BI) and Bias Harmfulness (BH), to characterize the impact landmarks exert on model reasoning, and establish a comprehensive benchmark, LandmarkBias-3K.
To mitigate landmark bias, they propose HoloGeo, an evidence-driven reasoning framework to improve the reliability of geo-localization. HoloGeo is supported by a high-quality dataset, BF-30k, annotated with structured multi-evidence bias-free reasoning chains. By incorporating multi-dimensional rewards, HoloGeo explicitly encourages balanced attention over diverse visual cues, enabling evidence-driven joint reasoning.
Extensive experiments show that HoloGeo not only maintains excellent performance on IM2GPS3K and YFCC4k, but also significantly outperforms existing open-source VLMs on LandmarkBias-3K, validating its effectiveness in robust geospatial reasoning.
Key Contributions
- Two quantitative bias metrics: Bias Intensity (BI) and Bias Harmfulness (BH)
- A new benchmark dataset: LandmarkBias-3K
- An evidence-driven reasoning framework: HoloGeo
- A high-quality training dataset: BF-30k with structured multi-evidence bias-free reasoning chains
- Multi-dimensional reward mechanism encouraging balanced use of visual evidence