Missing Values Handling for Machine Learning Portfolios
Обработка пропущенных значений для портфелей машинного обучения
2022-07-21
SCID: 54.1/g3tn6y3h
Discuss with AI
cross-sectional return predictorsexpectation-maximization (EM)imputation (cross-sectional mean)machine learning portfoliosmissing values handling
Figures from the paper
Abstract (AI)
We characterize the structure and origins of missingness for 159 cross-sectional return predictors and study missing value handling for portfolios constructed using machine learning. Simply imputing with cross-sectional means performs well compared to rigorous expectation-maximization methods. This stems from three facts about predictor data: (1) missingness occurs in large blocks organized by time, (2) cross-sectional correlations are small, and (3) missingness tends to occur in blocks organized by the underlying data source. As a result, observed data provide little information about missing data. Sophisticated imputations introduce estimation noise that can lead to underperformance if machine learning is not carefully applied.
Key Findings
1
For 159 cross-sectional return predictors, missingness structure and origins were characterized and analyzed for ML-constructed portfolios.
2
Observed data provide little information about the missing values given the blocky missingness and low cross-sectional correlation.
3
Simple cross-sectional mean imputation performs well compared to rigorous expectation-maximization methods for these predictors.
4
Sophisticated imputation methods can introduce estimation noise that causes underperformance unless machine learning is applied carefully.
5
Three data facts explain why simple imputation suffices: missingness occurs in large time-organized blocks; cross-sectional correlations are small; missingness aligns with underlying data-source blocks.
Research Object
Missing values in 159 cross-sectional return predictors used for machine-learning portfolio construction
Research Subject
Effectiveness of different missing-value handling/imputation methods (cross-sectional mean imputation vs expectation–maximization and sophisticated imputations) on machine-learning-constructed portfolio performance, given the structure and origins of missingness
Publication Details
Publication Date
2022-07-21
Journal
Publisher
ISSN
Cited by
20
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest