Unit 1: Data science and R fundamentals
Data Sciences Using R notes · PTU syllabus (MBA 962-18)
On this page
Unit summary
Data science combines statistics, computing and business knowledge to extract value from data, and R is one of its most widely used tools. This unit covers the components and roles in data science, big data, data pre-processing, supervised and unsupervised learning, business applications, installing R and its basic elements, R data interfaces, charts and graphs, and computing statistics in R.
After this unit you can
- Explain the components, roles and business applications of data science
- Explain big data, pre-processing and supervised and unsupervised learning
- Install R and use its basic elements and data interfaces
- Create charts and compute descriptive statistics in R
PTU syllabus topics
- Components and roles in data science
- big data/data pre-processing/supervised and unsupervised learning concepts
- business applications of data science
- introduction to R software installation and basic elements
- R data interfaces
- charts
- graphs and statistics
- mean/median/SD/variance/correlation/covariance through R
- Vector
- c(1, 2, 3)
- Data frame
- data.frame(x, y) or read.csv("file.csv")
- Summary
- summary(df), mean(), sd(), cor()
- Plot
- plot(), hist(), boxplot()
- Packages
- install.packages("ggplot2"); library(ggplot2)
Topic 1
Components of data science
Domain knowledge
Understanding the business problem
Mathematics and statistics
Probability, inference, modelling
Computer science
Programming, databases, algorithms
Data engineering
Collecting, storing and preparing data
Machine learning
Models that learn from data
Communication
Visualisation and storytelling
- 1
Business understanding
- 2
Data acquisition
- 3
Data preparation
- 4
Exploratory analysis
- 5
Modelling
- 6
Evaluation
- 7
Deployment and monitoring
Topic 2
Roles in data science
| Role | Main responsibility |
|---|---|
| Data scientist | Builds models, experiments and insights |
| Data analyst | Reports, dashboards, descriptive and diagnostic analysis |
| Data engineer | Builds data pipelines, warehouses and lakes |
| Machine learning engineer | Deploys and scales models in production |
| Business analyst or translator | Links business problems to analytics solutions |
| Data steward | Data quality, governance and privacy |
Topic 3
Big data and data pre-processing
- Big data: data sets too large or complex for traditional tools, described by the 5 Vs — volume, velocity, variety, veracity and value.
- Technologies: Hadoop (HDFS, MapReduce), Spark, NoSQL databases, cloud data platforms.
- Pre-processing: cleaning (missing values, outliers, duplicates), integration, transformation (scaling, encoding categories as dummies), reduction (feature selection, PCA), splitting into training and test sets.
Topic 4
Supervised and unsupervised learning
Data
Labelled — target variable known
Unlabelled
Goal
Predict or classify
Find structure or groups
Techniques
Regression, logistic regression, decision trees, SVM, random forests
Clustering (K-means, hierarchical, DBSCAN), PCA, association rules
Example
Predict loan default
Segment customers
- Reinforcement learning: an agent learns by trial and reward — used in recommendation, robotics and dynamic pricing.
Topic 5
Business applications of data science
- Marketing: segmentation, recommendation, churn prediction, marketing mix modelling.
- Finance: credit scoring, fraud detection, algorithmic trading, risk modelling.
- Operations: demand forecasting, inventory optimisation, predictive maintenance, route optimisation.
- HR: attrition prediction, resume screening, workforce planning.
- Healthcare and public sector: diagnosis support, disease surveillance, tax fraud analytics.
Topic 6
Introduction to R
- R is a free, open-source language and environment for statistical computing and graphics; RStudio (Posit) is the popular IDE with console, script editor, environment and plots panes.
- Installation: install R from CRAN, then RStudio; add packages with install.packages("name") and load with library(name).
- Popular packages: tidyverse (dplyr, ggplot2, readr, tidyr), caret and tidymodels (machine learning), randomForest, e1071 (SVM), cluster, rpart (decision trees).
Topic 7
Basic elements of R
Vector
c(10, 20, 30) — one data type
Matrix
matrix(1:6, nrow = 2) — two dimensions, one type
List
list(name = "A", marks = c(80, 90)) — mixed types
Data frame
data.frame(name, age) — table with columns of different types
Factor
factor(c("Low", "High")) — categorical data
- Data types: numeric, integer, character, logical (TRUE or FALSE), complex.
- Operators and assignment: x <- 5; arithmetic (+, −, the asterisk for multiplication, /, ^), relational (==, !=, >, <), logical (&, and the "or" operator).
- Control structures and functions: if–else, for and while loops, user-defined functions with function().
Topic 8
R data interfaces
| Task | R command |
|---|---|
| Read CSV | read.csv("sales.csv") or readr::read_csv() |
| Read Excel | readxl::read_excel("data.xlsx") |
| Write CSV | write.csv(df, "out.csv") |
| Connect to a database | DBI::dbConnect() with odbc or RMySQL |
| Read JSON or web data | jsonlite::fromJSON(url) |
| Inspect data | head(df), str(df), summary(df), dim(df) |
Topic 9
Charts and graphs in R
- Base graphics: plot(x, y), hist(x), boxplot(y ~ group), barplot(table(x)), pie(x).
- ggplot2: grammar of graphics — ggplot(df, aes(x = month, y = sales)) + geom_line(); layers for points, bars, smoothing, facets and themes.
Example
ggplot(sales, aes(x = region, y = revenue)) + geom_col() draws a bar chart of revenue by region in one line.
Topic 10
Statistics in R
Mean and median
mean(x), median(x)
Spread
sd(x), var(x), range(x), IQR(x), quantile(x)
Relationship
cor(x, y), cov(x, y)
Summary
summary(df) gives min, quartiles, mean and max
- Handling missing values: mean(x, na.rm = TRUE); is.na(); na.omit().
- Correlation matrix: cor(df[, c("price", "ads", "sales")]).
Key terms
- Data science
- Field extracting knowledge from data using statistics and computing
- Big data
- Data with high volume, velocity and variety
- Supervised learning
- Learning from labelled data
- Data frame
- R table with columns of different types
- ggplot2
- R package for layered graphics
Quick revision
- Components and life cycle of data science; roles.
- Big data 5 Vs; pre-processing steps.
- Supervised vs unsupervised vs reinforcement learning; applications.
- R and RStudio; packages; vectors, matrices, lists, data frames, factors.
- Data import and export; base graphics and ggplot2; mean, median, sd, var, cor, cov.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.Name the components of data science.
- Q2.What are the 5 Vs of big data?
- Q3.Distinguish supervised and unsupervised learning.
- Q4.What is a data frame in R?
- Q5.Write the R command to read a CSV file.
- Q6.Which R functions give variance and correlation?
Long-answer questions
- Q1.Explain the components, roles and life cycle of data science.
- Q2.Discuss big data, data pre-processing and types of machine learning.
- Q3.Explain the basic data structures and data interfaces of R.
- Q4.Explain how charts and descriptive statistics are produced in R.
Stuck on this unit?
Message SBS on WhatsApp for help with Data Sciences Using R, or to ask about studying MBA at Synetic.
