Missing Values Handling for Machine Learning Portfolios

Обработка пропущенных значений для портфелей машинного обучения
Andrew Y. Chen, Jack McCoy
2022-07-21

cross-sectional return predictorsexpectation-maximization (EM)imputation (cross-sectional mean)machine learning portfoliosmissing values handling
We characterize the structure and origins of missingness for 159 cross-sectional return predictors and study missing value handling for portfolios constructed using machine learning. Simply imputing with cross-sectional means performs well compared to rigorous expectation-maximization methods. This stems from three facts about predictor data: (1) missingness occurs in large blocks organized by time, (2) cross-sectional correlations are small, and (3) missingness tends to occur in blocks organized by the underlying data source. As a result, observed data provide little information about missing data. Sophisticated imputations introduce estimation noise that can lead to underperformance if machine learning is not carefully applied.
1
For 159 cross-sectional return predictors, missingness structure and origins were characterized and analyzed for ML-constructed portfolios.
2
Observed data provide little information about the missing values given the blocky missingness and low cross-sectional correlation.
3
Simple cross-sectional mean imputation performs well compared to rigorous expectation-maximization methods for these predictors.
4
Sophisticated imputation methods can introduce estimation noise that causes underperformance unless machine learning is applied carefully.
5
Three data facts explain why simple imputation suffices: missingness occurs in large time-organized blocks; cross-sectional correlations are small; missingness aligns with underlying data-source blocks.

Missing values in 159 cross-sectional return predictors used for machine-learning portfolio construction

Effectiveness of different missing-value handling/imputation methods (cross-sectional mean imputation vs expectation–maximization and sophisticated imputations) on machine-learning-constructed portfolio performance, given the structure and origins of missingness

Publication Details
Publication Date
2022-07-21
Journal
Publisher
ISSN
Cited by
20
Access Type
Author Information
Authors
Andrew Y. Chen
Jack McCoy
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%