Search Paper
  • Home
  • Login
  • Categories
  • Post URL
  • Academic Resources
  • Contact Us

 

Data Partitioning for Ensemble Model Building

google+
Views: 16                 

Author :  Ates Dagli1 , Niall McCarroll2 and Dmitry Vasilenko3

Affiliation :  1 IBM Big Data Analytics, 2 IBM Watson Machine Learning, 3 IBM Watson Cloud Platform and Data Science

Country :  USA

Category :  Cloud Computing

Volume, Issue, Month, Year :  7, 4, August, 2017

Abstract :


In distributed ensemble model-building algorithms, the performance and statistical validity of models are dependent on sizes of the input data partitions as well as the distribution of records among the partitions. Failure to correctly select and pre-process the data often results in the models which are not stable and do not perform well. This article introduces an optimized approach to building the ensemble models for very large data sets in distributed map-reduce environments using Pass-Stream-Merge (PSM) algorithm. To ensure the model correctness the input data is randomly distributed using the facilities built into mapreduce frameworks.

Keyword :  Ensemble Models, Pass-Stream-Merge, Big Data, Map-Reduce, Cloud

Journal/ Proceedings Name :  International Journal on Cloud Computing: Services and Architecture (IJCCSA)

URL :  https://aircconline.com/ijccsa/V7N4/7417ijccsa01.pdf

User Name : John Gails
Posted 22-07-2026 on 22:39:20 AEDT



Related Research Work

  • Lessons Learned From Implementing A Scalable Paas Service By Using Single Board Computers
  • Attribute Based Access Control (abac) For Ehr In Fog Computing Environment
  • Bridgegap: A Smart Behavioral Profiling And Matchmaking System For Bridge Players Using Real- Time Gameplay Analytics
  • A Health Research Collaboration Cloud Architecture

About Us | Post Cfp | Share URL Main | Share URL category | Post URL
All Rights Reserved @ Call for Papers - Conference & Journals