CatBoost for big data: an interdisciplinary review

CatBoost для больших данных: междисциплинарный обзор
Taghi M. Khoshgoftaar, John Hancock
2020-11-04

CatBoostGradient Boosted Decision Trees (GBDT)categorical heterogeneous dataclassification and regression taskshyper-parameter tuning
Gradient Boosted Decision Trees (GBDT's) are a powerful tool for classification and regression tasks in Big Data. Researchers should be familiar with the strengths and weaknesses of current implementations of GBDT's in order to use them effectively and make successful contributions. CatBoost is a member of the family of GBDT machine learning ensemble techniques. Since its debut in late 2018, researchers have successfully used CatBoost for machine learning studies involving Big Data. We take this opportunity to review recent research on CatBoost as it relates to Big Data, and learn best practices from studies that cast CatBoost in a positive light, as well as studies where CatBoost does not outshine other techniques, since we can learn lessons from both types of scenarios. Furthermore, as a Decision Tree based algorithm, CatBoost is well-suited to machine learning tasks involving categorical, heterogeneous data. Recent work across multiple disciplines illustrates CatBoost's effectiveness and shortcomings in classification and regression tasks. Another important issue we expose in literature on CatBoost is its sensitivity to hyper-parameters and the importance of hyper-parameter tuning. One contribution we make is to take an interdisciplinary approach to cover studies related to CatBoost in a single work. This provides researchers an in-depth understanding to help clarify proper application of CatBoost in solving problems. To the best of our knowledge, this is the first survey that studies all works related to CatBoost in a single publication.
1
CatBoost is an effective Gradient Boosted Decision Trees (GBDT) implementation for classification and regression tasks on Big Data, especially with categorical and heterogeneous features.
2
CatBoost is sensitive to hyper-parameters, making hyper-parameter tuning important for good performance.
3
The survey finds CatBoost has demonstrated effectiveness across multiple disciplines, but also exhibits shortcomings depending on task and dataset.
4
This work provides the first interdisciplinary survey consolidating studies on CatBoost, summarizing best practices and failure cases for proper application.

CatBoost algorithm applied to Big Data machine learning tasks

Performance, strengths and weaknesses, applicability, and hyper-parameter sensitivity/tuning of CatBoost for classification and regression on categorical and heterogeneous Big Data across disciplines

Publication Details
Publication Date
2020-11-04
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Taghi M. Khoshgoftaar
John Hancock
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%