graph-based methodslow-density separationsemi-supervised learningtransductionunlabeled data
Figures from the paper
Abstract (AI)
A comprehensive review of an area of machine learning that deals with the use of unlabeled data in classification problems: state-of-the-art algorithms, a taxonomy of the field, applications, benchmark experiments, and directions for future research. In the field of machine learning, semi-supervised learning (SSL) occupies the middle ground, between supervised learning (in which all training examples are labeled) and unsupervised learning (in which no label data are given). Interest in SSL has increased in recent years, particularly because of application domains in which unlabeled data are plentiful, such as images, text, and bioinformatics. This first comprehensive overview of SSL presents state-of-the-art algorithms, a taxonomy of the field, selected applications, benchmark experiments, and perspectives on ongoing and future research.Semi-Supervised Learning first presents the key assumptions and ideas underlying the field: smoothness, cluster or low-density separation, manifold structure, and transduction. The core of the book is the presentation of SSL methods, organized according to algorithmic strategies. After an examination of generative models, the book describes algorithms that implement the low-density separation assumption, graph-based methods, and algorithms that perform two-step learning. The book then discusses SSL applications and offers guidelines for SSL practitioners by analyzing the results of extensive benchmark experiments. Finally, the book looks at interesting directions for SSL research. The book closes with a discussion of the relationship between semi-supervised learning and transduction.
Key Findings
1
Applications in images, text, and bioinformatics motivate SSL because these domains provide abundant unlabeled data.
2
Extensive benchmark experiments support practical guidelines, while the review identifies ongoing and future research directions and clarifies SSL’s relationship to transduction.
3
SSL methods are taxonomized by algorithmic strategies, including generative models, low-density-separation methods, graph-based approaches, and two-step learning.
4
Semi-supervised learning uses unlabeled data alongside labeled examples, bridging supervised and unsupervised learning for classification.
5
The field is organized around smoothness, cluster or low-density separation, manifold structure, and transduction assumptions.
Research Object
Semi-supervised learning methods and algorithms for classification using both labeled and unlabeled data
Research Subject
state-of-the-art algorithms, underlying assumptions, taxonomy, applications, benchmark performance, and future research directions
Publication Details
Publication Date
2006-09-22
Journal
Publisher
ISSN
Cited by
4337
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai6
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs1997
Statistical Learning Theory1999
Neural Networks for Pattern Recognition1995
Additive logistic regression: a statistical view of boosting (With discussion and a rejoinder by the authors)2000
Pattern Recognition and Neural Networks1996
Neural Networks and the Bias/Variance Dilemma1992
Cited by7
The Graph Neural Network Model2008
A survey on semi-supervised learning2019
Toward Causal Representation Learning2021
Machine learning & artificial intelligence in the quantum domain: a review of recent progress2018
A Unifying Review of Deep and Shallow Anomaly Detection2021
Advancements and Challenges in Machine Learning: A Comprehensive Review of Models, Libraries, Applications, and Algorithms2023
Potential, challenges and future directions for deep learning in prognostics and health management applications2020