Sehaa: A Big Data Analytics Tool for Healthcare Symptoms and Diseases Detection Using Twitter, Apache Spark, and Machine Learning

Sehaa: инструмент анализа больших данных для выявления симптомов и заболеваний в сфере здравоохранения с использованием Twitter, Apache Spark и машинного обучения
Aiiad Albeshri, Rashid Mehmood, Omer Rana, Iyad Katib, Shoayee Alotaibi
2020-02-19

Apache SparkArabic text classificationNaive BayesTwitter datahealthcare disease detection
Smartness, which underpins smart cities and societies, is defined by our ability to engage with our environments, analyze them, and make decisions, all in a timely manner. Healthcare is the prime candidate needing the transformative capability of this smartness. Social media could enable a ubiquitous and continuous engagement between healthcare stakeholders, leading to better public health. Current works are limited in their scope, functionality, and scalability. This paper proposes Sehaa, a big data analytics tool for healthcare in the Kingdom of Saudi Arabia (KSA) using Twitter data in Arabic. Sehaa uses Naive Bayes, Logistic Regression, and multiple feature extraction methods to detect various diseases in the KSA. Sehaa found that the top five diseases in Saudi Arabia in terms of the actual afflicted cases are dermal diseases, heart diseases, hypertension, cancer, and diabetes. Riyadh and Jeddah need to do more in creating awareness about the top diseases. Taif is the healthiest city in the KSA in terms of the detected diseases and awareness activities. Sehaa is developed over Apache Spark allowing true scalability. The dataset used comprises 18.9 million tweets collected from November 2018 to September 2019. The results are evaluated using well-known numerical criteria (Accuracy and F1-Score) and are validated against externally available statistics.
1
Riyadh and Jeddah require greater awareness efforts regarding the leading diseases, while Taif is identified as the healthiest city based on detected diseases and awareness activities.
2
Sehaa is a scalable Apache Spark tool that analyzes Arabic Twitter data to detect healthcare symptoms and diseases in Saudi Arabia.
3
The five most prevalent detected diseases in Saudi Arabia are dermal diseases, heart diseases, hypertension, cancer, and diabetes.
4
The system is evaluated with Accuracy and F1-Score and validated against externally available health statistics.
5
Using 18.9 million tweets collected from November 2018 to September 2019, Sehaa applies Naive Bayes, Logistic Regression, and multiple feature-extraction methods.

Arabic Twitter healthcare data from Saudi Arabia (KSA), comprising 18.9 million tweets about symptoms and diseases

Detection and spatial analysis of diseases and public-health awareness patterns, including disease prevalence, city-level differences, and detection performance

Publication Details
Publication Date
2020-02-19
Journal
Publisher
ISSN
Cited by
121
Access Type
Author Information
Authors
Aiiad Albeshri
Rashid Mehmood
Omer Rana
Iyad Katib
Shoayee Alotaibi
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%