Big Data Analytics
Subject Overview
The fourth Elective-II option, covering big data fundamentals and high-performance architecture (HDFS, MapReduce, YARN), advanced analytical methods (clustering, decision trees, Naive Bayes, association rules, recommendation systems), stream computing and real-time analytics, and Hadoop implementation and deployment. A 4-credit elective theory paper.
Unit-wise Syllabus
4 units — click WhatsApp below to get the full notes for each
Unit 1: Big data fundamentals
Evolution and characteristics of big data, best practices, big data use cases, understanding big data storage, high-performance architecture overview — HDFS, MapReduce and YARN, MapReduce programming model
Unit 2: Advanced analytical methods
K-means clustering use cases and diagnostics, decision tree algorithms and evaluation, Naive Bayes classification, association rules and the Apriori algorithm, evaluation of candidate rules, finding association and similarity, collaborative/content-based/knowledge-based/hybrid recommendation systems
Unit 3: Stream computing and real-time analytics
Stream data model and architecture, sampling and filtering streams, counting distinct elements, estimating moments, decaying windows, real-time analytics platform applications, real-time sentiment analysis, stock market prediction, graph analytics for big data
Unit 4: Hadoop implementation
Hadoop cluster components and architecture, Hadoop ecosystem, evaluation criteria for distributed MapReduce runtimes, enterprise-grade Hadoop deployment and implementation
