Paper Overview
Field: Computer Vision Authors: Yejun Zhang, Xinjue Wang, Zihan Wang Published: 2026-07-04 arXiv: 2507.03228
Abstract
Descriptor-free visual localization eliminates high-dimensional descriptor storage, preserves scene privacy, and simplifies map maintenance, yet its accuracy still lags far behind descriptor-based pipelines. The paper identifies this gap as stemming from insufficient geometric discriminability in geometry-only matching. Without visual appearance, current methods underutilize local geometric cues, lack global context among keypoints, and overfit to a single keypoint detector.
The authors further observe that descriptor-free matching naturally enables multi-detector training, since heterogeneous keypoints can be optimized in a shared geometry-only space without aligning descriptor spaces.
The GeoMix Framework
GeoMix is a descriptor-free 2D-3D matching framework that strengthens geometric discriminability at three levels:
- Local level: directional and distance-aware embeddings enrich neighborhood aggregation with fine-grained spatial structure;
- Global level: learnable context nodes aggregate and redistribute scene-level information via cross-attention, resolving ambiguities beyond local receptive fields;
- Training level: Mix-Training leverages a detector-agnostic geometric space to learn representations across multiple keypoint detectors.
- 75th-percentile rotation error reduced by 89%;
- Translation error reduced by up to 90%;
- Zero-shot generalization to unseen detectors;
- Significantly narrowed gap with descriptor-based pipelines.
Results
Extensive experiments on MegaDepth, Cambridge Landmarks, 7Scenes, and Aachen Day-Night show that GeoMix achieves new state-of-the-art results among descriptor-free methods: