Unit 3: Regression and risk analytics
Financial Analytics notes · PTU syllabus (MBA 965-26)
On this page
Unit summary
Regression, classification and clustering answer core questions in investments, credit and fraud. This unit covers regression and factor models in finance, forecasting financial variables, logistic regression for credit scoring, ROC curves and the confusion matrix, Benford's Law for fraud detection, and K-means clustering for customer segmentation.
After this unit you can
- Apply regression and factor models in finance
- Build credit scoring models with logistic regression and evaluate them
- Apply Benford's Law to detect fraud
- Segment customers with K-means clustering
PTU syllabus topics
- Regression models and factor models in finance
- forecasting financial variables
- logistic regression for credit scoring
- ROC curves and confusion matrix
- Benford's Law for fraud detection
- K-means clustering for customer segmentation
Credit scoring
Logistic regression
Probability of default
Model quality
ROC curve, confusion matrix
AUC, accuracy
Fraud detection
Benford's Law
Suspicious digit patterns
Customer segmentation
K-means clustering
Groups of similar clients
Topic 1
Regression and factor models in finance
Market model (CAPM regression)
Ri − Rf = α + β (Rm − Rf) + e
Fama–French three-factor
Ri − Rf = α + β (Rm − Rf) + s SMB + h HML + e
Interpretation
β measures market risk; α is abnormal return
- Uses: estimating beta and cost of equity, performance attribution, risk models, building factor portfolios.
Topic 2
Forecasting financial variables
- Regression on drivers (sales on GDP and prices), time series models for interest rates and exchange rates, machine learning models with many predictors; evaluate out-of-sample and beware overfitting.
Topic 3
Logistic regression for credit scoring
- Use: binary outcomes — churn or not, default or not, buy or not.
Model
glm(churn ~ tenure + complaints + plan, family = binomial, data = train)
Probabilities
predict(model, test, type = "response")
Odds ratio
exp(coef(model))
- Classify using a cut-off (e.g., 0.5) and evaluate with a confusion matrix and AUC.
- Credit scorecards: variables such as income, repayment history, bureau score (CIBIL), loan-to-income, employment; scores map to probability of default (PD) for pricing and approval.
Topic 4
ROC curves and the confusion matrix
| Predicted positive | Predicted negative | |
|---|---|---|
| Actual positive | True positive (TP) | False negative (FN) |
| Actual negative | False positive (FP) | True negative (TN) |
Accuracy
(TP + TN) ÷ total
Precision
TP ÷ (TP + FP)
Recall (sensitivity)
TP ÷ (TP + FN)
Specificity
TN ÷ (TN + FP)
F1 score
2 × precision × recall ÷ (precision + recall)
Example
Of 1,000 customers, 100 churn. The model flags 120: 80 actual churners (TP), 40 non-churners (FP). Precision = 80 ÷ 120 = 67%; recall = 80 ÷ 100 = 80%; accuracy = (80 + 860) ÷ 1,000 = 94%.
- ROC curve and AUC: trade-off between sensitivity and false positive rate across cut-offs (pROC package).
- Imbalanced data: accuracy misleads — use precision, recall, F1, AUC; rebalance with oversampling (SMOTE) or class weights.
Topic 5
Benford's Law for fraud detection
- Benford's Law: in many naturally occurring data sets, the leading digit d appears with probability log10(1 + 1/d) — 1 appears about 30.1% of the time, 9 about 4.6%.
Digit 1
30.1%
Digit 2
17.6%
Digit 3
12.5%
Digit 5
7.9%
Digit 9
4.6%
- Use: auditors compare actual first-digit frequencies of invoices, expenses or claims with Benford's distribution (chi-square test); large deviations flag possible manipulation.
- Limits: not applicable to assigned numbers, data with fixed ranges or thresholds.
Example
An expense audit finds an unusual excess of claims starting with 4 and 9 just below approval limits of ₹5,000 and ₹10,000 — a red flag for splitting or fabricating claims.
Topic 6
K-means clustering for customer segmentation
- Variables: balances, transactions, product holdings, channel use, age, income, credit usage.
- Process: standardise variables, choose k (elbow, silhouette), run K-means, profile clusters, design strategies for each.
Example
A bank finds four clusters — young digital spenders, affluent investors, salaried savers and small-business owners — and tailors products and channels to each.
Key terms
- Beta
- Sensitivity of a stock's return to market return
- Alpha
- Return above that explained by risk factors
- Probability of default
- Likelihood a borrower defaults
- Benford's Law
- Distribution of leading digits in natural data
- Customer segmentation
- Grouping customers with similar characteristics
Quick revision
- Market model; Fama–French; beta and alpha.
- Forecasting financial variables; out-of-sample testing.
- Logistic regression for PD; scorecards.
- Confusion matrix, precision, recall; ROC–AUC.
- Benford's Law frequencies and uses; K-means segmentation.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.What does beta measure?
- Q2.State the Fama–French three factors.
- Q3.Why is logistic regression used in credit scoring?
- Q4.What does AUC measure?
- Q5.State Benford's probability for the digit 1.
- Q6.What is K-means clustering?
Long-answer questions
- Q1.Explain regression and factor models in finance.
- Q2.Explain credit scoring using logistic regression and its evaluation.
- Q3.Explain Benford's Law and its use in fraud detection.
- Q4.Explain customer segmentation using K-means clustering.
Stuck on this unit?
Message SBS on WhatsApp for help with Financial Analytics, or to ask about studying MBA at Synetic.
