Educational data mining: prediction of students' academic performance using machine learning algorithms
Интеллектуальный анализ образовательных данных: прогнозирование академической успеваемости студентов с использованием алгоритмов машинного обучения
2022-03-03
SCID: 54.1/d6mz98gc
Discuss with AI
educational data miningmachine learning algorithmsrandom forestsstudent academic performance predictionsupport vector machines
Figures from the paper
Abstract (AI)
Abstract Educational data mining has become an effective tool for exploring the hidden relationships in educational data and predicting students' academic achievements. This study proposes a new model based on machine learning algorithms to predict the final exam grades of undergraduate students, taking their midterm exam grades as the source data. The performances of the random forests, nearest neighbour, support vector machines, logistic regression, Naïve Bayes, and k-nearest neighbour algorithms, which are among the machine learning algorithms, were calculated and compared to predict the final exam grades of the students. The dataset consisted of the academic achievement grades of 1854 students who took the Turkish Language-I course in a state University in Turkey during the fall semester of 2019–2020. The results show that the proposed model achieved a classification accuracy of 70–75%. The predictions were made using only three types of parameters; midterm exam grades, Department data and Faculty data. Such data-driven studies are very important in terms of establishing a learning analysis framework in higher education and contributing to the decision-making processes. Finally, this study presents a contribution to the early prediction of students at high risk of failure and determines the most effective machine learning methods.
Key Findings
1
A machine-learning model was developed to predict undergraduate students' final exam grades using midterm grades as source data.
2
Random Forests, Nearest Neighbour, Support Vector Machines, Logistic Regression, Naïve Bayes, and k-Nearest Neighbour were evaluated and compared for prediction performance.
3
The dataset comprised academic grades of 1,854 students from a Turkish Language-I course at a Turkish state university (fall 2019–2020).
4
The study demonstrates the model's utility for early prediction of students at high risk of failure and identifies the most effective machine learning methods for this task.
5
Using only midterm exam grades, department, and faculty as features, the proposed model achieved classification accuracy of 70–75% on the dataset.
Research Object
Undergraduate students who took the Turkish Language-I course (academic records: midterm grades, department, faculty) at a state university (fall 2019–2020)
Research Subject
Prediction of final exam grades / early identification of students at high risk of failure using machine learning classification algorithms (random forests, k-NN, SVM, logistic regression, Naïve Bayes) based on midterm grades, department, and faculty
Publication Details
Publication Date
2022-03-03
Journal
Publisher
ISSN
Cited by
687
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest