Unit 2 of 4 · MBA Sem 3

Unit 2: Probability and regression

Data Sciences Using R notes · PTU syllabus (MBA 962-18)

3 min read5 topics10 exam questions
On this page
  1. Unit summary
  2. Probability for data science
  3. Linear and multiple regression
  4. Logistic regression
  5. Decision trees
  6. Support vector machines
  7. Key terms
  8. Quick revision
  9. Important questions

Unit summary

Probability underpins prediction, and regression and classification models are the workhorses of applied data science. This unit covers probability theory and Bayes' theorem, linear, multiple and logistic regression, decision trees and support vector machines, with their implementation in R.

After this unit you can

  • Apply probability and Bayes' theorem in data science
  • Build linear and multiple regression models in R
  • Build logistic regression models in R
  • Build decision trees and SVMs in R

PTU syllabus topics

  • Probability theory for data science (Bayes theorem)
  • linear/multiple/logistic regression
  • decision tree and Support Vector Machine (SVM)
Key termsModels in R
Linear regression
lm(y ~ x1 + x2, data = df)
Logistic regression
glm(y ~ x, family = binomial)
Decision tree
rpart(y ~ ., data = df)
SVM
svm(y ~ ., data = df) from e1071
Predict
predict(model, newdata)
1

Topic 1

Probability for data science

  • Probability: likelihood of an event, between 0 and 1; key rules — addition, multiplication, conditional probability.
  • Distributions: binomial (dbinom, pbinom), Poisson (dpois), normal (dnorm, pnorm, qnorm) — used in R for modelling counts, rates and continuous variables.
Key formulasBayes' theorem
  • Bayes

    P(A given B) = P(B given A) × P(A) ÷ P(B)

  • Total probability

    P(B) = Σ P(B given Ai) × P(Ai)

Example

2% of transactions are fraudulent. A model flags 90% of frauds and 5% of genuine transactions. P(fraud given flagged) = 0.9 × 0.02 ÷ (0.9 × 0.02 + 0.05 × 0.98) = 0.018 ÷ 0.067 ≈ 27%.

  • Naive Bayes classifier: e1071::naiveBayes(target ~ ., data = train).
2

Topic 2

Linear and multiple regression

Key formulasRegression in R
  • Simple linear

    model <- lm(sales ~ ads, data = df)

  • Multiple

    model <- lm(sales ~ ads + price + outlets, data = df)

  • Output

    summary(model) — coefficients, p-values, R², adjusted R², F-statistic

  • Prediction

    predict(model, newdata = test)

  • Diagnostics: plot(model) for residuals; car::vif(model) for multicollinearity; check linearity, normality and constant variance.

Example

Coefficient for ads = 2.4 (p below 0.01) means each extra ₹1 lakh of advertising is associated with ₹2.4 lakh more sales, holding price and outlets constant.

3

Topic 3

Logistic regression

  • Use: binary outcomes — churn or not, default or not, buy or not.
Key formulasLogistic regression in R
  • Model

    glm(churn ~ tenure + complaints + plan, family = binomial, data = train)

  • Probabilities

    predict(model, test, type = "response")

  • Odds ratio

    exp(coef(model))

  • Classify using a cut-off (e.g., 0.5) and evaluate with a confusion matrix and AUC.
4

Topic 4

Decision trees

  • rpart package: tree <- rpart(default ~ income + age + emi_ratio, data = train, method = "class"); rpart.plot(tree) to visualise.
  • Splitting criteria: Gini index (default in rpart) or information gain; pruning with the complexity parameter (cp) to avoid overfitting.
  • Strengths: interpretable rules, handle mixed data; weakness: unstable — small data changes can change the tree (addressed by ensembles).
5

Topic 5

Support vector machines

  • Idea: find the hyperplane with the maximum margin between classes; kernels (linear, polynomial, radial) handle non-linear boundaries.
  • In R: e1071::svm(class ~ ., data = train, kernel = "radial", cost = 1); tune cost and gamma with tune().
  • Use: text classification, image recognition, credit scoring; accurate but less interpretable.

Key terms

Conditional probability
Probability of an event given another has occurred
lm()
R function for linear regression
glm()
R function for generalised linear models, including logistic regression
Complexity parameter
rpart setting controlling tree pruning
Kernel
Function mapping data to higher dimensions in SVM

Quick revision

  • Probability rules; distributions in R; Bayes' theorem; naive Bayes.
  • lm() for simple and multiple regression; summary; diagnostics; VIF.
  • glm(family = binomial) for logistic regression; odds ratios; cut-off.
  • rpart trees; Gini; pruning with cp.
  • SVM: margin, kernels, e1071::svm, tuning.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.State Bayes' theorem.
  2. Q2.Write the R command for a multiple regression of sales on price and ads.
  3. Q3.What does adjusted R² measure?
  4. Q4.Which family is used in glm() for logistic regression?
  5. Q5.What is pruning in decision trees?
  6. Q6.What is the role of a kernel in SVM?

Long-answer questions

  1. Q1.Explain probability theory and Bayes' theorem with a data science application.
  2. Q2.Explain linear and multiple regression and their implementation in R.
  3. Q3.Explain logistic regression and its interpretation.
  4. Q4.Compare decision trees and support vector machines.

Stuck on this unit?

Message SBS on WhatsApp for help with Data Sciences Using R, or to ask about studying MBA at Synetic.

WhatsApp us