A survey on semi-supervised learning

Обзор обучения с полуучителем
Holger H. Hoos, Jesper E. van Engelen
2019-11-15

neural network-based modelssemi-supervised classificationsemi-supervised clustering assumptionsemi-supervised learningunlabelled data
Abstract Semi-supervised learning is the branch of machine learning concerned with using labelled as well as unlabelled data to perform certain learning tasks. Conceptually situated between supervised and unsupervised learning, it permits harnessing the large amounts of unlabelled data available in many use cases in combination with typically smaller sets of labelled data. In recent years, research in this area has followed the general trends observed in machine learning, with much attention directed at neural network-based models and generative learning. The literature on the topic has also expanded in volume and scope, now encompassing a broad spectrum of theory, algorithms and applications. However, no recent surveys exist to collect and organize this knowledge, impeding the ability of researchers and engineers alike to utilize it. Filling this void, we present an up-to-date overview of semi-supervised learning methods, covering earlier work as well as more recent advances. We focus primarily on semi-supervised classification, where the large majority of semi-supervised learning research takes place. Our survey aims to provide researchers and practitioners new to the field as well as more advanced readers with a solid understanding of the main approaches and algorithms developed over the past two decades, with an emphasis on the most prominent and currently relevant work. Furthermore, we propose a new taxonomy of semi-supervised classification algorithms, which sheds light on the different conceptual and methodological approaches for incorporating unlabelled data into the training process. Lastly, we show how the fundamental assumptions underlying most semi-supervised learning algorithms are closely connected to each other, and how they relate to the well-known semi-supervised clustering assumption.
1
It focuses primarily on semi-supervised classification, where most research in the field has been conducted, and organizes prominent methods from the past two decades.
2
It shows that the fundamental assumptions underlying many semi-supervised learning algorithms are closely interconnected and related to the established semi-supervised clustering assumption.
3
The paper introduces a new taxonomy of semi-supervised classification algorithms based on conceptual and methodological strategies for incorporating unlabelled data.
4
The survey addresses the lack of recent comprehensive reviews, aiming to improve accessibility of semi-supervised learning knowledge for researchers and practitioners.
5
The survey provides an up-to-date overview of semi-supervised learning methods, covering foundational work and recent advances across theory, algorithms, and applications.

semi-supervised learning methods, primarily semi-supervised classification algorithms

the conceptual and methodological approaches, fundamental assumptions, and taxonomy for incorporating unlabelled data into the training process

Publication Details
Publication Date
2019-11-15
Journal
Publisher
ISSN
Cited by
2647
Access Type
Author Information
Authors
Holger H. Hoos
Jesper E. van Engelen
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%