A survey of open source tools for machine learning with big data in the Hadoop ecosystem
Обзор инструментов с открытым исходным кодом для машинного обучения на больших данных в экосистеме Hadoop
2015-11-05
SCID: 54.1/rpdktm7e
Discuss with AI
Hadoop ecosystembig data machine learningdistributed processingmachine learning frameworksscalability
Figures from the paper
Abstract (AI)
With an ever-increasing amount of options, the task of selecting machine learning tools for big data can be difficult. The available tools have advantages and drawbacks, and many have overlapping uses. The world’s data is growing rapidly, and traditional tools for machine learning are becoming insufficient as we move towards distributed and real-time processing. This paper is intended to aid the researcher or professional who understands machine learning but is inexperienced with big data. In order to evaluate tools, one should have a thorough understanding of what to look for. To that end, this paper provides a list of criteria for making selections along with an analysis of the advantages and drawbacks of each. We do this by starting from the beginning, and looking at what exactly the term “big data” means. From there, we go on to the Hadoop ecosystem for a look at many of the projects that are part of a typical machine learning architecture and an understanding of how everything might fit together. We discuss the advantages and disadvantages of three different processing paradigms along with a comparison of engines that implement them, including MapReduce, Spark, Flink, Storm, and H 2 O. We then look at machine learning libraries and frameworks including Mahout, MLlib, SAMOA, and evaluate them based on criteria such as scalability, ease of use, and extensibility. There is no single toolkit that truly embodies a one-size-fits-all solution, so this paper aims to help make decisions smoother by providing as much information as possible and quantifying what the tradeoffs will be. Additionally, throughout this paper, we review recent research in the field using these tools and talk about possible future directions for toolkit-based learning.
Key Findings
1
It defines selection criteria for evaluating tools, including scalability, ease of use, extensibility, and tradeoffs among advantages and drawbacks.
2
It evaluates machine-learning libraries and frameworks such as Mahout, MLlib, and SAMOA within broader Hadoop-based architectures.
3
The paper concludes that no single toolkit provides a universal solution, making workload-specific tradeoff analysis necessary for tool selection.
4
The paper surveys open-source machine-learning tools for big data within the Hadoop ecosystem, targeting readers inexperienced with big-data systems.
5
The survey compares processing paradigms and engines, including MapReduce, Spark, Flink, Storm, and H2O, for distributed and real-time machine learning.
Research Object
Open-source machine-learning tools and frameworks for big-data processing in the Hadoop ecosystem
Research Subject
Their comparative capabilities and trade-offs, particularly scalability, ease of use, extensibility, processing paradigms, and suitability for distributed and real-time machine learning
Publication Details
Publication Date
2015-11-05
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest