Summary
Researchers present a machine-learning framework to cross-match X-ray sources from the Chandra Source Catalog (CSC v2.1) with optical sources from Gaia Data Release 3. Unlike purely spatial cross-matching, the method uses source properties such as magnitudes, colors, and distances to identify true counterparts, detect chance coincidences, and resolve ambiguities when multiple plausible candidates exist. A training set of high-confidence matches is built with NWAY, a Bayesian cross-matching framework, and a gradient-boosted classifier (LightGBM) is trained on features from both catalogs. Of about 254k unique X-ray sources, counterparts are found for about 113k, including roughly 7k with plausible multiple candidates. About 20k sources that separation-based matching associates are left unmatched, with half attributed to chance coincidences. Validation on the Chandra Orion Ultradeep Project (COUP) shows the machine-learning matches reproduce 95% of NWAY cross-matches without using positional information. The team releases catalogs of ~113k counterparts, ~7k alternative matches, and ~20k ambiguous associations (arXiv:2506.14975).
Overview
arXiv: 2506.14975
Authors: V. Samuel Pérez-Díaz, Vinay L. Kashyap, Joshua D. Ingram
This post introduces a framework to cross-match sources from the Chandra Source Catalog (CSC v2.1) with optical sources from Gaia Data Release 3. Unlike purely spatial approaches, it uses source properties such as magnitudes, colors, and distances to identify true counterparts, detect chance coincidences, and resolve ambiguities when multiple plausible candidates exist.
Method
- A training set of high-confidence matches is defined using NWAY, a Bayesian cross-matching framework that accounts for positional errors and source densities.
- A gradient-boosted classifier (LightGBM) is trained on a variety of features from both catalogs.
Key results
- Of the ~254k unique X-ray sources, counterparts are found for ~113k sources, of which plausible multiple counterparts exist for ~7k.
- ~20k sources for which separation-based cross-matching finds a match receive no counterpart here; half of these are attributed to chance coincidences.
- Validation on the Chandra Orion Ultradeep Project (COUP) shows the machine-learning matches reproduce 95% of NWAY cross-matches without using any positional information.
Data release
The authors release:
- A catalog of the ~113k Chandra-Gaia counterparts
- ~7k alternative matches
- ~20k ambiguous NWAY associations
These support future population studies of sources detectable by both Chandra and Gaia. The paper also discusses limitations and provides a generalization of the framework applicable to other cross-matching scenarios.
---
*Originally collected on 2026-06-19.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177981511