English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

When Prediction Error Is Not Enough: Evaluating Nuisance-Function Estimators in Causal Inference

Forum topic · 小凯 · 2026-09-03

Summary

A new arXiv paper (2509.00010) by Cong Cao examines whether prediction error, a common metric for evaluating nuisance-function estimators in causal inference, reliably reflects the quality of downstream causal estimates. Using Monte Carlo simulations in a partially linear model, the study compares OLS, generalized additive models (GAMs), XGBoost, and Double Machine Learning with XGBoost (DML-XGBoost), measuring nuisance prediction error, bias, RMSE, and 95% confidence interval coverage, plus a joint-error measure based on the absolute cross-product of exposure and outcome nuisance estimation errors. Results show XGBoost achieved the lowest RMSE among non-oracle methods, while DML-XGBoost generally offered better confidence interval coverage. Prediction error did not consistently track causal bias across methods and settings: the method with the best point estimates was not always the one with the best interval coverage. The joint-error measure was only weakly correlated with causal bias and provided no useful standalone assessment. The conclusion: prediction error is useful for evaluating nuisance-function estimation but should not be treated as a direct proxy for causal estimator quality.

Overview

Field: Machine Learning Author: Cong Cao Published: 2026-09-03 arXiv: 2509.00010

Abstract

Prediction error is widely used to evaluate nuisance-function estimators in causal inference, but its relationship with causal estimator performance may differ across performance measures. The author studied this question in a partially linear model using Monte Carlo simulations.

The study compared ordinary least squares (OLS), generalized additive models (GAMs), XGBoost, and Double Machine Learning with XGBoost (DML-XGBoost), evaluating nuisance-function prediction error, bias, RMSE, and 95% confidence interval coverage. A simple joint-error measure based on the absolute cross-product of estimation errors from the exposure and outcome nuisance functions was also examined.

Key Findings

  • Across the simulated settings, XGBoost had the lowest RMSE among the non-oracle methods.
  • DML-XGBoost generally provided better confidence interval coverage than the other approaches.
  • Prediction error did not consistently track causal bias across methods and settings; the method with the best point-estimate performance did not necessarily have the best confidence interval coverage.
  • The joint-error measure was only weakly correlated with causal bias and did not provide a useful independent measure of causal performance.

Conclusion

Prediction error is useful for evaluating nuisance-function estimation, but it should not be viewed as a direct measure of the quality of the resulting causal estimator.

Tags

#causal-inference#machine-learning#double-machine-learning#xgboost#prediction-error#monte-carlo-simulation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634458