Unit 3: Predictive modeling
Business Analytics for Decision Making notes · PTU syllabus (MBA 201-26)
On this page
Unit summary
Predictive models estimate what is likely to happen so that managers can act early. This unit covers business forecasting methods, correlation and regression, testing regression assumptions — multicollinearity, heteroscedasticity and autocorrelation — linear and logistic regression, K-nearest neighbours, decision trees and model evaluation using SPSS.
After this unit you can
- Explain business forecasting methods
- Apply correlation and linear regression and test their assumptions
- Apply logistic regression, KNN and decision trees
- Evaluate predictive models
PTU syllabus topics
- Business forecasting methods
- correlation and regression analysis
- testing assumptions (multicollinearity, heteroscedasticity, autocorrelation)
- linear regression
- logistic regression
- K-nearest neighbors
- decision trees
- model evaluation techniques using SPSS
Linear regression
A number
Next quarter's sales
Logistic regression
Probability of a class
Will this customer churn?
k-NN
Class from similar cases
Credit risk band
Decision tree
Class or number by rules
Loan approval
Topic 1
Business forecasting methods
Qualitative
Expert opinion, Delphi method, sales force composite, market surveys
Time series
Moving averages, exponential smoothing, trend projection, ARIMA
Causal (econometric)
Regression on drivers such as price, income, advertising
Machine learning
Decision trees, ensembles, neural networks
- Steps: define the purpose and horizon, collect data, choose the method, build and validate, forecast, monitor errors.
- Forecast accuracy: MAD (mean absolute deviation), MSE, MAPE (mean absolute percentage error).
Topic 2
Correlation analysis
r = Σ(x − x̄)(y − ȳ) / √[Σ(x − x̄)² × Σ(y − ȳ)²] r lies between −1 and +1: +1 perfect positive, −1 perfect negative, 0 no linear correlation.
Example
x: 1, 2, 3, 4, 5 and y: 2, 4, 5, 4, 5. With x̄ = 3 and ȳ = 4: Σdxdy = 6, Σdx² = 10, Σdy² = 6, so r = 6/√60 ≈ 0.77 — fairly high positive correlation.
The probable error PE = 0.6745 (1 − r²)/√n; if r > 6 PE, correlation is significant.
Topic 3
Linear regression
Regression estimates the value of one variable from another. Two regression lines exist:
- y on x: y − ȳ = byx (x − x̄), with byx = r × σy/σx
- x on y: x − x̄ = bxy (y − ȳ), with bxy = r × σx/σy
The principle of least squares fits the line that minimises the sum of squared vertical deviations.
- Geometric mean
- r = ±√(byx × bxy)
- Same sign
- byx, bxy and r all have the same sign
- Both cannot exceed 1
- If one is above 1, the other is below 1
- Independent of origin
- But not of scale
- Lines meet
- At (x̄, ȳ)
Purpose
Measures strength of relationship
Predicts one variable from another
Symmetry
rxy = ryx
byx ≠ bxy
Cause and effect
Not implied
Treats one as dependent
- Multiple regression: Y = b0 + b1X1 + b2X2 + … + e; each b shows the effect of one variable holding the others constant.
- Goodness of fit: R² (share of variation explained), adjusted R² (penalises extra variables), F-test for the overall model, t-tests for individual coefficients.
Exam tip
In SPSS: Analyze → Regression → Linear; read Model Summary (R²), ANOVA (F, Sig.) and Coefficients (B, Beta, t, Sig.).
Topic 4
Testing the assumptions of regression
Linearity
Scatter and residual plots
Normality of residuals
Histogram, P–P plot of residuals
No multicollinearity
VIF below 10 (stricter: below 5), tolerance above 0.1
Homoscedasticity
Residual vs predicted plot shows no funnel; Breusch–Pagan test
No autocorrelation
Durbin–Watson statistic near 2 (1.5 to 2.5 acceptable)
- Remedies: drop or combine correlated predictors (multicollinearity), transform variables or use robust errors (heteroscedasticity), add lagged variables or use time-series models (autocorrelation).
Exam tip
In SPSS: under Linear Regression → Statistics tick Collinearity diagnostics and Durbin–Watson; under Plots put ZRESID against ZPRED.
Topic 5
Logistic regression
Logistic regression predicts a binary outcome (buy or not, default or not) by modelling the log-odds as a linear function of predictors.
Probability
P = 1 ÷ (1 + e^−(b0 + b1X1 + …))
Odds ratio
Exp(B) — change in odds for a one-unit rise in X
- Output: omnibus test, Nagelkerke R², Hosmer–Lemeshow test, classification table, Exp(B).
Exam tip
In SPSS: Analyze → Regression → Binary Logistic.
Example
A bank's model gives Exp(B) = 1.8 for "missed EMI in last year" — a borrower with a missed EMI has 1.8 times the odds of default.
Topic 6
K-nearest neighbours and decision trees
- KNN: classifies a new case by the majority class of its k closest cases (distance on standardised variables); simple but slow on large data and sensitive to the choice of k. SPSS: Analyze → Classify → Nearest Neighbor.
- Decision tree: splits data repeatedly on the variable that best separates outcomes, producing if–then rules. Algorithms: CHAID, CART (CRT), QUEST, C4.5. SPSS: Analyze → Classify → Tree.
Income below ₹30,000
Check credit score — below 650 high risk; otherwise medium risk
Income ₹30,000 or more
Check existing EMIs — above 50% of income medium risk; otherwise low risk
- Strengths of trees: easy to explain, handle mixed data; weakness: may overfit — controlled by pruning and minimum node size.
Topic 7
Model evaluation techniques
- Validation: split data into training and test sets (e.g., 70:30), or use k-fold cross-validation, to check performance on unseen data.
R² and adjusted R²
Variation explained (regression)
RMSE
√(mean of squared errors)
Accuracy
(TP + TN) ÷ total
Precision
TP ÷ (TP + FP)
Recall (sensitivity)
TP ÷ (TP + FN)
F1 score
2 × precision × recall ÷ (precision + recall)
- ROC curve and AUC: plot sensitivity against 1 − specificity; AUC near 1 is excellent, 0.5 is no better than chance.
- Overfitting vs underfitting: a model too complex fits noise; one too simple misses patterns.
Key terms
- Multicollinearity
- High correlation among predictors
- Heteroscedasticity
- Unequal variance of residuals
- Durbin–Watson
- Statistic testing autocorrelation of residuals
- Logistic regression
- Model for a binary outcome
- Confusion matrix
- Table of predicted vs actual classes
Quick revision
- Forecasting: qualitative, time series, causal, machine learning; MAD, MAPE.
- Correlation; linear and multiple regression; R², adjusted R².
- Assumptions: linearity, normality, VIF, homoscedasticity, Durbin–Watson.
- Logistic regression (Exp(B)); KNN; decision trees (CHAID, CART).
- Evaluation: train–test, cross-validation, confusion matrix, ROC–AUC.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.Name two qualitative forecasting methods.
- Q2.What is multicollinearity and how is it detected?
- Q3.What does a Durbin–Watson value of 2 indicate?
- Q4.When is logistic regression used?
- Q5.What is the k in KNN?
- Q6.Define precision and recall.
Long-answer questions
- Q1.Discuss the methods of business forecasting.
- Q2.Explain linear regression and the tests of its assumptions.
- Q3.Explain logistic regression, KNN and decision trees as predictive tools.
- Q4.Explain the techniques used to evaluate predictive models.
Stuck on this unit?
Message SBS on WhatsApp for help with Business Analytics for Decision Making, or to ask about studying MBA at Synetic.
