Deep learning applications and challenges in big data analytics

Приложения и проблемы глубокого обучения в аналитике больших данных
Taghi M. Khoshgoftaar, Maryam M. Najafabadi, Flavio Villanustre, Naeem Seliya, Randall Wald, Edin Muharemagic
2015-02-23

Big Data AnalyticsDeep Learningdistributed computingsemantic indexingstreaming data
Abstract Big Data Analytics and Deep Learning are two high-focus of data science. Big Data has become important as many organizations both public and private have been collecting massive amounts of domain-specific information, which can contain useful information about problems such as national intelligence, cyber security, fraud detection, marketing, and medical informatics. Companies such as Google and Microsoft are analyzing large volumes of data for business analysis and decisions, impacting existing and future technology. Deep Learning algorithms extract high-level, complex abstractions as data representations through a hierarchical learning process. Complex abstractions are learnt at a given level based on relatively simpler abstractions formulated in the preceding level in the hierarchy. A key benefit of Deep Learning is the analysis and learning of massive amounts of unsupervised data, making it a valuable tool for Big Data Analytics where raw data is largely unlabeled and un-categorized. In the present study, we explore how Deep Learning can be utilized for addressing some important problems in Big Data Analytics, including extracting complex patterns from massive volumes of data, semantic indexing, data tagging, fast information retrieval, and simplifying discriminative tasks. We also investigate some aspects of Deep Learning research that need further exploration to incorporate specific challenges introduced by Big Data Analytics, including streaming data, high-dimensional data, scalability of models, and distributed computing. We conclude by presenting insights into relevant future works by posing some questions, including defining data sampling criteria, domain adaptation modeling, defining criteria for obtaining useful data abstractions, improving semantic indexing, semi-supervised learning, and active learning.
1
Challenges remain for applying Deep Learning to Big Data, notably handling streaming data, high-dimensional data, model scalability, and distributed computing.
2
Deep Learning can address Big Data tasks such as extracting complex patterns, semantic indexing, data tagging, fast information retrieval, and simplifying discriminative tasks.
3
Deep Learning is well-suited for Big Data Analytics because it can learn high-level complex abstractions from largely unlabeled, massive datasets via hierarchical learning.
4
Further research directions include defining data sampling criteria, domain adaptation modeling, criteria for useful data abstractions, improving semantic indexing, semi-supervised learning, and active learning.

Deep learning methods applied to Big Data analytics

Use and challenges of deep learning for extracting complex patterns, semantic indexing, data tagging, fast information retrieval, simplifying discriminative tasks, and handling streaming, high-dimensional, scalable and distributed big-data scenarios

Publication Details
Publication Date
2015-02-23
Journal
Publisher
ISSN
Cited by
2612
Access Type
Author Information
Authors
Taghi M. Khoshgoftaar
Maryam M. Najafabadi
Flavio Villanustre
Naeem Seliya
Randall Wald
Edin Muharemagic
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%