Data Sciences Using R
Subject Overview
The second Business Analytics group elective, covering the components and business applications of data science with R software basics, probability theory and regression/classification techniques, ensemble methods and clustering, and evaluation methods for data mining results. A 4-credit elective theory paper.
Unit-wise Syllabus
4 units — click WhatsApp below to get the full notes for each
Unit 1: Data science and R fundamentals
Components and roles in data science, big data/data pre-processing/supervised and unsupervised learning concepts, business applications of data science, introduction to R software installation and basic elements, R data interfaces, charts, graphs and statistics, mean/median/SD/variance/correlation/covariance through R
Unit 2: Probability and regression
Probability theory for data science (Bayes theorem), linear/multiple/logistic regression, decision tree and Support Vector Machine (SVM)
Unit 3: Ensemble methods and clustering
Bagging, random forests, boosting, K-means clustering, K-medoids, agglomerative and hierarchical clustering, X-means, DBSCAN
Unit 4: Evaluation and validation
Methods for estimating classifier performance — cross-validation, holdout method, bootstrap method, confusion matrix, assessing statistical significance of data mining results, advanced topics (scalable ML, big data techniques, stream data mining, social networks)
