Two-point correlation function estimation with contaminated data
Phys. Rev. D 113, 123522 – Published 11 June, 2026
DOI: https://doi.org/10.1103/nhr2-7w6x
Abstract
The two-point correlation function (2PCF) is a cornerstone of precision cosmology, yet its estimation from imaging surveys is vulnerable to contamination and incompleteness arising from imperfect target selection and pipeline-level inclusion decisions. In practice, the scientific target is a physically defined population (e.g., galaxies in a redshift or luminosity range), while the working catalog is constructed from noisy measurements and selection cuts, leading to mismatches between true and observed inclusion. These errors are rarely spatially uniform: they correlate with survey depth, observing conditions, and foreground structure, and can imprint spurious large-scale power or suppress the true clustering signal. High-resolution, high-SNR spectroscopic samples provide gold-standard inclusion in the target population, but are typically available for only a small subset of objects. We introduce a prediction-powered Landy-Szalay (PP-LS) estimator that combines noisy inclusion labels over the full catalog with exact labels on a small spectroscopic subset, while preserving the standard random-catalog normalization that corrects for survey geometry and selection. PP-LS debiases pair counts through residual-based, design-weighted correction terms computed only on the labeled subset, requiring no probability calibration, no known misclassification rates or class priors, no explicit spatial model of contamination, nor forward modeling of the systematics. Under simple random sampling of the labeled subset, we establish recovery of the oracle (true-label) Landy-Szalay pair counts and hence consistency for the target 2PCF. In controlled simulations with clustered and spatially structured contaminants, PP-LS removes the bias of naïve catalog-level estimators while achieving substantially lower variance than spectroscopic-only clustering. The result is a statistically principled, computationally lightweight estimator that integrates directly with standard pair-counting pipelines and supports robust clustering inference in next-generation surveys.