Big Five Personality Trait Prediction Based on User Comments

Прогнозирование черт личности по модели Большой пятёрки на основе пользовательских комментариев
Fumito Masui, Michał Ptaszyński, Kit-May Shum
2025-05-20

BERT (base and large)Big Five personality trait predictionPANDORA dataset (Reddit)R2RMSERoBERTa (RoBERTa large)single-model vs multiple-model approachtrait intercorrelationstransformer-based language models
The study of personalities is a major component of human psychology, and with an understanding of personality traits, practical applications can be used in various domains, such as mental health care, predicting job performance, and optimising marketing strategies. This study explores the prediction of Big Five personality trait scores from online comments using transformer-based language models, focusing on improving the model performance with a larger dataset and investigating the role of intercorrelations among traits. Using the PANDORA dataset from Reddit, the RoBERTa and BERT models, including both the base and large variants, were fine-tuned and evaluated to determine their effectiveness in personality trait prediction. Compared to previous work, our study utilises a significantly larger dataset to enhance the model’s generalisation and robustness. The results indicate that RoBERTa outperforms BERT across most metrics, with RoBERTa large achieving the best overall performance. In addition to evaluating the overall predictive accuracy, this study investigates the impact of intercorrelations among personality traits. A comparative analysis is conducted between a single-model approach, which predicts all five traits simultaneously, and a multiple-model approach, fine-tuning the models independently and each predicting a single trait. The findings reveal that the single-model approach achieves a lower RMSE and higher R2 values, highlighting the importance of incorporating trait intercorrelations in improving the prediction accuracy. Furthermore, RoBERTa large demonstrated a stronger ability to capture these intercorrelations compared to previous studies. These findings emphasise the potential of transformer-based models in personality computing and underscore the importance of leveraging both larger datasets and intercorrelations to enhance predictive performance.
1
A single-model approach that predicts all five traits simultaneously yields lower RMSE and higher R2 than separate single-trait models.
2
RoBERTa large more effectively captures intercorrelations among personality traits than models used in previous studies.
3
RoBERTa outperforms BERT across most evaluation metrics, with RoBERTa large achieving the best overall performance.
4
Transformer-based language models (RoBERTa and BERT) can predict Big Five personality trait scores from Reddit comments.
5
Using a significantly larger PANDORA dataset improves model generalisation and robustness compared to previous work.

Online user comments from Reddit (PANDORA dataset)

Prediction of Big Five personality trait scores from those comments using transformer-based language models, including effects of dataset size and trait intercorrelations on model performance

Publication Details
Publication Date
2025-05-20
Journal
Publisher
ISSN
Cited by
7
Access Type
Author Information
Authors
Fumito Masui
Michał Ptaszyński
Kit-May Shum
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%