EaSFE: Scalable and Efficient Feature Engineering for Boosting Machine Learning Performance

EaSFE: Масштабируемая и эффективная инженерия признаков для повышения производительности машинного обучения
Jian Chen, Yile Chen, Zheng Zhenya, Zeyi Wen, Yawen Chen, Jin Huang
2026-05-14

EaSFEchunk-based data processinghigh-dimensional sparse data handlingparallel and distributed executionscalable feature engineering
Feature engineering plays a critical role in machine learning (ML), but existing methods often struggle with high computational cost and limited scalability when applied to large-scale and sparse datasets. In this article, we propose EaSFE, an efficient and scalable feature engineering framework that unifies feature generation, filtering, and evaluation in an end-to-end manner. EaSFE is designed to efficiently construct and select informative features while explicitly considering computational and memory constraints. To achieve scalability, EaSFE incorporates parallel and distributed execution mechanisms, as well as a chunk-based data processing strategy that enables memory-efficient feature engineering on large datasets. In addition, EaSFE adopts tailored storage and execution strategies to handle high-dimensional sparse data effectively. Extensive experiments on multiple real-world datasets demonstrate that EaSFE consistently improves predictive performance (e.g., 5% accuracy improvement in poker ) while substantially enhancing efficiency (i.e., over 10x speedup) compared to existing feature engineering methods. In addition, EaSFE is demonstrated to scale to large and sparse datasets, successfully handling datasets with over 119 million training instances and 54 million features.
1
EaSFE achieves substantial efficiency gains, reporting over 10x speedup compared to existing feature engineering methods.
2
EaSFE employs tailored storage and execution strategies to effectively handle high-dimensional sparse data.
3
EaSFE explicitly considers computational and memory constraints when constructing and selecting informative features.
4
EaSFE is an end-to-end feature engineering framework that unifies feature generation, filtering, and evaluation for efficiency and scalability.
5
EaSFE scales to very large datasets, demonstrated on data with over 119 million training instances and 54 million features.
6
EaSFE uses parallel and distributed execution plus a chunk-based data processing strategy to enable memory-efficient feature engineering on large datasets.
7
EaSFE yields consistent predictive improvements (e.g., 5% accuracy improvement on the poker dataset) over existing feature engineering methods.

EaSFE feature engineering framework for large-scale and sparse datasets

Efficient and scalable feature engineering including unified feature generation, filtering, evaluation, memory- and computation-aware selection, parallel/distributed and chunk-based processing to boost ML predictive performance and efficiency

Publication Details
Publication Date
2026-05-14
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Jian Chen
Yile Chen
Zheng Zhenya
Zeyi Wen
Yawen Chen
Jin Huang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%