On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation

О переобучении при выборе модели и последующей смещённости выбора в оценке качества
Gavin C. Cawley, Nicola L. C. Talbot
2010-03-01

cross-validation based model selectionk-fold cross-validationover-fitting in model selectionselection bias in performance evaluationvariance of model selection criterion
Model selection strategies for machine learning algorithms typically involve the numerical opti-misation of an appropriate model selection criterion, often based on an estimator of generalisation performance, such as k-fold cross-validation. The error of such an estimator can be broken down into bias and variance components. While unbiasedness is often cited as a beneficial quality of a model selection criterion, we demonstrate that a low variance is at least as important, as a non-negligible variance introduces the potential for over-fitting in model selection as well as in training the model. While this observation is in hindsight perhaps rather obvious, the degradation in perfor-mance due to over-fitting the model selection criterion can be surprisingly large, an observation that appears to have received little attention in the machine learning literature to date. In this paper, we show that the effects of this form of over-fitting are often of comparable magnitude to differences in performance between learning algorithms, and thus cannot be ignored in empirical evaluation. Furthermore, we show that some common performance evaluation practices are susceptible to a form of selection bias as a result of this form of over-fitting and hence are unreliable. We dis-cuss methods to avoid over-fitting in model selection and subsequent selection bias in performance evaluation, which we hope will be incorporated into best practice. While this study concentrates on cross-validation based model selection, the findings are quite general and apply to any model selection practice involving the optimisation of a model selection criterion evaluated over a finite sample of data, including maximisation of the Bayesian evidence and optimisation of performance bounds.
1
Common performance evaluation practices can suffer selection bias due to over-fitting of the model selection criterion and thus be unreliable.
2
Findings, while demonstrated for cross-validation, generalize to any model selection that optimizes a criterion on a finite data sample (e.g., Bayesian evidence, performance bounds).
3
High variance in model selection criteria can cause over-fitting during model selection and training, even if the criterion is unbiased.
4
Methods to reduce over-fitting in model selection and subsequent selection bias are necessary and should be adopted as best practice.
5
Performance degradation from over-fitting the model selection criterion can be surprisingly large, comparable to algorithm performance differences.

Model selection process for machine learning algorithms (especially cross-validation–based selection)

Over-fitting of the model selection criterion and resulting selection bias in performance evaluation, including effects of estimator variance and methods to avoid these biases

Publication Details
Publication Date
2010-03-01
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Gavin C. Cawley
Nicola L. C. Talbot
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%