← back

Optimizing the extraction of information from redshift probability distribution functions

📄 arXiv:2607.26822 · 📥 PDF · 2026-07-29 · astro-ph.CO

Authors: Rodrigo Duarte [arXiv · scholar] , Valerio Marra [arXiv · scholar]

🕰 Orloj analysis

7.9
Total score
8.5
Consistency
8.0
Quality
AD relevance

Tento článek představuje turboPDZ, framework strojového učení pro optimalizaci extrakce bodových odhadů a měr spolehlivosti z distribučních funkcí pravděpodobnosti (PDZ) fotometrických rudých posuvů. Framework zlepšuje přesnost a spolehlivost oproti stávajícím metodám a je veřejně dostupný.

💡 Jde o solidní metodologický příspěvek, který přináší významné zlepšení v přesnosti a spolehlivosti fotometrických rudých posuvů pro analýzy velkoškálové struktury.

Categories: INF-7 COS-1 MET-2 MET-1

✓ code_available, falsifiable, modest_claims, limit_reductions

📄 Abstract

Photometric redshifts are essential for large-scale structure analyses, yet extracting optimal point estimates and reliability measures from the probability distribution functions (PDZs) delivered by photo-$z$ pipelines remains an open challenge. We introduce turboPDZ, a machine-learning framework that optimizes both quantities directly from the PDZ. We apply the framework to PDZs from the three independent HSC-SSP PDR3 pipelines (DEmP, DNNz, Mizuki) across Wide and DUD layers. Each PDZ is compressed via PCA and combined with summary descriptors; a multilayer perceptron, optimized with Optuna under a composite objective, produces the optimized point estimate $z_{\rm ml}$. A second network, trained in log-space and calibrated, yields the uncertainty $σ_{\rm ml}$, from which the reliability score $r_{\rm ml}$ is derived via percentile ranking. $z_{\rm ml}$ outperforms the catalog $z_{\rm best}$ in $σ_{\rm NMAD}$ and $η_{0.15}$ across all six pipeline-layer combinations. $r_{\rm ml}$ filters galaxies more efficiently than the catalog risk and confidence indicators, as measured by the area under the $σ_{\rm NMAD}$ and $η_{0.15}$ versus retained-fraction curves. For Mizuki, the template-fitting pipeline, the catalog indicators fail dramatically, with AUC values up to ten times larger than those of $r_{\rm ml}$, whereas $r_{\rm ml}$ correctly identifies unreliable objects across all redshift regimes. Feature-importance analysis reveals complementary patterns: point estimation is dominated by PCA components and location statistics, while reliability estimation depends on PCA components and peak statistics. The pipeline is survey-independent, publicly available at https://github.com/valerio-marra/turboPDZ, and trained models plus optimized quantities are released as a value-added catalog.

📄 arXiv abstract page 📥 PDF