Additive logistic regression: a statistical view of boosting (With discussion and a rejoinder by the authors)

Аддитивная логистическая регрессия: статистический взгляд на бустинг (с обсуждением и ответом авторов)
Jerome H. Friedman, Robert Tibshirani, Trevor Hastie
2000-04-01

Additive logistic regressionAdditive modelingBoostingMaximum Bernoulli likelihoodMultinomial likelihood
Boosting is one of the most important recent developments in classification methodology. Boosting works by sequentially applying a classification algorithm to reweighted versions of the training data and then taking a weighted majority vote of the sequence of classifiers thus produced. For many classification algorithms, this simple strategy results in dramatic improvements in performance. We show that this seemingly mysterious phenomenon can be understood in terms of well-known statistical principles, namely additive modeling and maximum likelihood. For the two-class problem, boosting can be viewed as an approximation to additive modeling on the logistic scale using maximum Bernoulli likelihood as a criterion. We develop more direct approximations and show that they exhibit nearly identical results to boosting. Direct multiclass generalizations based on multinomial likelihood are derived that exhibit performance comparable to other recently proposed multiclass generalizations of boosting in most situations, and far superior in some. We suggest a minor modification to boosting that can reduce computation, often by factors of 10 to 50. Finally, we apply these insights to produce an alternative formulation of boosting decision trees. This approach, based on best-first truncated tree induction, often leads to better performance, and can provide interpretable descriptions of the aggregate decision rule. It is also much faster computationally, making it more suitable to large-scale data mining applications.
1
A minor modification to boosting can reduce computation by factors of 10 to 50.
2
An alternative formulation of boosting decision trees using best-first truncated tree induction often yields better performance, faster computation, and more interpretable aggregate decision rules suitable for large-scale data mining.
3
Boosting can be interpreted as additive modeling on the logistic scale using maximum Bernoulli likelihood for two-class problems.
4
Direct approximations to boosting derived from this statistical view produce nearly identical results to traditional boosting.
5
Multiclass generalizations based on multinomial likelihood achieve performance comparable to recent multiclass boosting methods and are superior in some situations.

Boosting algorithms for classification (including AdaBoost and decision-tree boosting)

Statistical interpretation and approximation of boosting as additive logistic regression using (Bernoulli/multinomial) maximum likelihood, extensions to multiclass problems, computational modifications, and an alternative best-first truncated tree formulation improving performance and speed

Publication Details
Publication Date
2000-04-01
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Jerome H. Friedman
Robert Tibshirani
Trevor Hastie
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%