Overview
Field: Machine Learning Author: Cong Cao Published: 2026-09-03 arXiv: 2509.00010
Abstract
Prediction error is widely used to evaluate nuisance-function estimators in causal inference, but its relationship with causal estimator performance may differ across performance measures. The author studied this question in a partially linear model using Monte Carlo simulations.
The study compared ordinary least squares (OLS), generalized additive models (GAMs), XGBoost, and Double Machine Learning with XGBoost (DML-XGBoost), evaluating nuisance-function prediction error, bias, RMSE, and 95% confidence interval coverage. A simple joint-error measure based on the absolute cross-product of estimation errors from the exposure and outcome nuisance functions was also examined.
Key Findings
- Across the simulated settings, XGBoost had the lowest RMSE among the non-oracle methods.
- DML-XGBoost generally provided better confidence interval coverage than the other approaches.
- Prediction error did not consistently track causal bias across methods and settings; the method with the best point-estimate performance did not necessarily have the best confidence interval coverage.
- The joint-error measure was only weakly correlated with causal bias and did not provide a useful independent measure of causal performance.
Conclusion
Prediction error is useful for evaluating nuisance-function estimation, but it should not be viewed as a direct measure of the quality of the resulting causal estimator.