EaSFE: Scalable and Efficient Feature Engineering for Boosting Machine Learning Performance
EaSFE: Масштабируемая и эффективная инженерия признаков для повышения производительности машинного обучения
2026-05-14
SCID: 54.1/2rsvajn9
Discuss with AI
EaSFEchunk-based data processinghigh-dimensional sparse data handlingparallel and distributed executionscalable feature engineering
Figures from the paper
Abstract (AI)
Feature engineering plays a critical role in machine learning (ML), but existing methods often struggle with high computational cost and limited scalability when applied to large-scale and sparse datasets. In this article, we propose EaSFE, an efficient and scalable feature engineering framework that unifies feature generation, filtering, and evaluation in an end-to-end manner. EaSFE is designed to efficiently construct and select informative features while explicitly considering computational and memory constraints. To achieve scalability, EaSFE incorporates parallel and distributed execution mechanisms, as well as a chunk-based data processing strategy that enables memory-efficient feature engineering on large datasets. In addition, EaSFE adopts tailored storage and execution strategies to handle high-dimensional sparse data effectively. Extensive experiments on multiple real-world datasets demonstrate that EaSFE consistently improves predictive performance (e.g., 5% accuracy improvement in poker ) while substantially enhancing efficiency (i.e., over 10x speedup) compared to existing feature engineering methods. In addition, EaSFE is demonstrated to scale to large and sparse datasets, successfully handling datasets with over 119 million training instances and 54 million features.
Key Findings
1
EaSFE achieves substantial efficiency gains, reporting over 10x speedup compared to existing feature engineering methods.
2
EaSFE employs tailored storage and execution strategies to effectively handle high-dimensional sparse data.
3
EaSFE explicitly considers computational and memory constraints when constructing and selecting informative features.
4
EaSFE is an end-to-end feature engineering framework that unifies feature generation, filtering, and evaluation for efficiency and scalability.
5
EaSFE scales to very large datasets, demonstrated on data with over 119 million training instances and 54 million features.
6
EaSFE uses parallel and distributed execution plus a chunk-based data processing strategy to enable memory-efficient feature engineering on large datasets.
7
EaSFE yields consistent predictive improvements (e.g., 5% accuracy improvement on the poker dataset) over existing feature engineering methods.
Research Object
EaSFE feature engineering framework for large-scale and sparse datasets
Research Subject
Efficient and scalable feature engineering including unified feature generation, filtering, evaluation, memory- and computation-aware selection, parallel/distributed and chunk-based processing to boost ML predictive performance and efficiency
Publication Details
Publication Date
2026-05-14
Journal
Publisher
ISSN
Cited by
0
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest