English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training

Forum topic · 小凯 · 2026-07-05

Summary

GeoMix (arXiv:2507.03228) is a descriptor-free 2D-3D matching framework for visual localization developed by Yejun Zhang, Xinjue Wang, and Zihan Wang. Descriptor-free localization avoids storing high-dimensional descriptors, preserves scene privacy, and simplifies map maintenance, but has trailed descriptor-based pipelines in accuracy. The authors attribute this gap to insufficient geometric discriminability in geometry-only matching: local geometry cues are underused, keypoints lack global context, and models overfit to a single keypoint detector. GeoMix addresses this at three levels: directional and distance-aware embeddings for fine-grained local neighborhood aggregation; learnable context nodes that aggregate and redistribute scene-level information via cross-attention; and Mix-Training, which exploits the detector-agnostic geometric space to train across multiple keypoint detectors without descriptor alignment. Experiments on MegaDepth, Cambridge Landmarks, 7Scenes, and Aachen Day-Night set a new state of the art among descriptor-free methods, reducing 75th-percentile rotation error by 89% and translation error by up to 90%, generalizing zero-shot to unseen detectors, and narrowing the gap with descriptor-based pipelines.

Paper Overview

  • Field: Computer Vision
  • Authors: Yejun Zhang, Xinjue Wang, Zihan Wang
  • arXiv: 2507.03228
  • Abstract

    Descriptor-free visual localization eliminates high-dimensional descriptor storage, preserves scene privacy, and simplifies map maintenance, yet its accuracy still lags far behind descriptor-based pipelines. The authors identify this gap as stemming from insufficient geometric discriminability in geometry-only matching: without visual appearance, current methods underutilize local geometry cues, lack global context among keypoints, and overfit to a single keypoint detector.

    A key observation is that descriptor-free matching naturally enables multi-detector training, since heterogeneous keypoints can be optimized in a shared geometry-only space without aligning descriptor spaces.

    The GeoMix Framework

    GeoMix strengthens geometric discriminability at three levels:

  • Local: directional and distance-aware embeddings enrich neighborhood aggregation with fine-grained spatial structure.
  • Global: learnable context nodes aggregate and redistribute scene-level information via cross-attention, resolving ambiguities beyond local receptive fields.
  • Training: Mix-Training leverages the detector-agnostic geometric space to learn representations across multiple keypoint detectors.
  • Results

    Extensive experiments on MegaDepth, Cambridge Landmarks, 7Scenes, and Aachen Day-Night show that GeoMix achieves a new state of the art among descriptor-free methods:

  • 75th-percentile rotation error reduced by 89%
  • Translation error reduced by up to 90%
  • Zero-shot generalization to unseen detectors
  • Significantly narrowed gap with descriptor-based pipelines
---

*Source: arXiv 2507.03228*

Tags

#visual-localization#descriptor-free#2d-3d-matching#computer-vision#keypoint-detection#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208427