Unit 2: Probability and regression
Data Sciences Using R notes · PTU syllabus (MBA 962-18)
On this page
Unit summary
Probability underpins prediction, and regression and classification models are the workhorses of applied data science. This unit covers probability theory and Bayes' theorem, linear, multiple and logistic regression, decision trees and support vector machines, with their implementation in R.
After this unit you can
- Apply probability and Bayes' theorem in data science
- Build linear and multiple regression models in R
- Build logistic regression models in R
- Build decision trees and SVMs in R
PTU syllabus topics
- Probability theory for data science (Bayes theorem)
- linear/multiple/logistic regression
- decision tree and Support Vector Machine (SVM)
- Linear regression
- lm(y ~ x1 + x2, data = df)
- Logistic regression
- glm(y ~ x, family = binomial)
- Decision tree
- rpart(y ~ ., data = df)
- SVM
- svm(y ~ ., data = df) from e1071
- Predict
- predict(model, newdata)
Topic 1
Probability for data science
- Probability: likelihood of an event, between 0 and 1; key rules — addition, multiplication, conditional probability.
- Distributions: binomial (dbinom, pbinom), Poisson (dpois), normal (dnorm, pnorm, qnorm) — used in R for modelling counts, rates and continuous variables.
Bayes
P(A given B) = P(B given A) × P(A) ÷ P(B)
Total probability
P(B) = Σ P(B given Ai) × P(Ai)
Example
2% of transactions are fraudulent. A model flags 90% of frauds and 5% of genuine transactions. P(fraud given flagged) = 0.9 × 0.02 ÷ (0.9 × 0.02 + 0.05 × 0.98) = 0.018 ÷ 0.067 ≈ 27%.
- Naive Bayes classifier: e1071::naiveBayes(target ~ ., data = train).
Topic 2
Linear and multiple regression
Simple linear
model <- lm(sales ~ ads, data = df)
Multiple
model <- lm(sales ~ ads + price + outlets, data = df)
Output
summary(model) — coefficients, p-values, R², adjusted R², F-statistic
Prediction
predict(model, newdata = test)
- Diagnostics: plot(model) for residuals; car::vif(model) for multicollinearity; check linearity, normality and constant variance.
Example
Coefficient for ads = 2.4 (p below 0.01) means each extra ₹1 lakh of advertising is associated with ₹2.4 lakh more sales, holding price and outlets constant.
Topic 3
Logistic regression
- Use: binary outcomes — churn or not, default or not, buy or not.
Model
glm(churn ~ tenure + complaints + plan, family = binomial, data = train)
Probabilities
predict(model, test, type = "response")
Odds ratio
exp(coef(model))
- Classify using a cut-off (e.g., 0.5) and evaluate with a confusion matrix and AUC.
Topic 4
Decision trees
- rpart package: tree <- rpart(default ~ income + age + emi_ratio, data = train, method = "class"); rpart.plot(tree) to visualise.
- Splitting criteria: Gini index (default in rpart) or information gain; pruning with the complexity parameter (cp) to avoid overfitting.
- Strengths: interpretable rules, handle mixed data; weakness: unstable — small data changes can change the tree (addressed by ensembles).
Topic 5
Support vector machines
- Idea: find the hyperplane with the maximum margin between classes; kernels (linear, polynomial, radial) handle non-linear boundaries.
- In R: e1071::svm(class ~ ., data = train, kernel = "radial", cost = 1); tune cost and gamma with tune().
- Use: text classification, image recognition, credit scoring; accurate but less interpretable.
Key terms
- Conditional probability
- Probability of an event given another has occurred
- lm()
- R function for linear regression
- glm()
- R function for generalised linear models, including logistic regression
- Complexity parameter
- rpart setting controlling tree pruning
- Kernel
- Function mapping data to higher dimensions in SVM
Quick revision
- Probability rules; distributions in R; Bayes' theorem; naive Bayes.
- lm() for simple and multiple regression; summary; diagnostics; VIF.
- glm(family = binomial) for logistic regression; odds ratios; cut-off.
- rpart trees; Gini; pruning with cp.
- SVM: margin, kernels, e1071::svm, tuning.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.State Bayes' theorem.
- Q2.Write the R command for a multiple regression of sales on price and ads.
- Q3.What does adjusted R² measure?
- Q4.Which family is used in glm() for logistic regression?
- Q5.What is pruning in decision trees?
- Q6.What is the role of a kernel in SVM?
Long-answer questions
- Q1.Explain probability theory and Bayes' theorem with a data science application.
- Q2.Explain linear and multiple regression and their implementation in R.
- Q3.Explain logistic regression and its interpretation.
- Q4.Compare decision trees and support vector machines.
Stuck on this unit?
Message SBS on WhatsApp for help with Data Sciences Using R, or to ask about studying MBA at Synetic.
