Unit 3 of 4 · MBA Sem 4

Unit 3: Regression and risk analytics

Financial Analytics notes · PTU syllabus (MBA 965-26)

3 min read6 topics10 exam questions
On this page
  1. Unit summary
  2. Regression and factor models in finance
  3. Forecasting financial variables
  4. Logistic regression for credit scoring
  5. ROC curves and the confusion matrix
  6. Benford's Law for fraud detection
  7. K-means clustering for customer segmentation
  8. Key terms
  9. Quick revision
  10. Important questions

Unit summary

Regression, classification and clustering answer core questions in investments, credit and fraud. This unit covers regression and factor models in finance, forecasting financial variables, logistic regression for credit scoring, ROC curves and the confusion matrix, Benford's Law for fraud detection, and K-means clustering for customer segmentation.

After this unit you can

  • Apply regression and factor models in finance
  • Build credit scoring models with logistic regression and evaluate them
  • Apply Benford's Law to detect fraud
  • Segment customers with K-means clustering

PTU syllabus topics

  • Regression models and factor models in finance
  • forecasting financial variables
  • logistic regression for credit scoring
  • ROC curves and confusion matrix
  • Benford's Law for fraud detection
  • K-means clustering for customer segmentation
ComparisonAnalytics in finance applications
Technique
Output

Credit scoring

Logistic regression

Probability of default

Model quality

ROC curve, confusion matrix

AUC, accuracy

Fraud detection

Benford's Law

Suspicious digit patterns

Customer segmentation

K-means clustering

Groups of similar clients

1

Topic 1

Regression and factor models in finance

Key formulasFactor models
  • Market model (CAPM regression)

    Ri − Rf = α + β (Rm − Rf) + e

  • Fama–French three-factor

    Ri − Rf = α + β (Rm − Rf) + s SMB + h HML + e

  • Interpretation

    β measures market risk; α is abnormal return

  • Uses: estimating beta and cost of equity, performance attribution, risk models, building factor portfolios.
2

Topic 2

Forecasting financial variables

  • Regression on drivers (sales on GDP and prices), time series models for interest rates and exchange rates, machine learning models with many predictors; evaluate out-of-sample and beware overfitting.
3

Topic 3

Logistic regression for credit scoring

  • Use: binary outcomes — churn or not, default or not, buy or not.
Key formulasLogistic regression in R
  • Model

    glm(churn ~ tenure + complaints + plan, family = binomial, data = train)

  • Probabilities

    predict(model, test, type = "response")

  • Odds ratio

    exp(coef(model))

  • Classify using a cut-off (e.g., 0.5) and evaluate with a confusion matrix and AUC.
  • Credit scorecards: variables such as income, repayment history, bureau score (CIBIL), loan-to-income, employment; scores map to probability of default (PD) for pricing and approval.
4

Topic 4

ROC curves and the confusion matrix

Predicted positivePredicted negative
Actual positiveTrue positive (TP)False negative (FN)
Actual negativeFalse positive (FP)True negative (TN)
Key formulasClassifier metrics
  • Accuracy

    (TP + TN) ÷ total

  • Precision

    TP ÷ (TP + FP)

  • Recall (sensitivity)

    TP ÷ (TP + FN)

  • Specificity

    TN ÷ (TN + FP)

  • F1 score

    2 × precision × recall ÷ (precision + recall)

Example

Of 1,000 customers, 100 churn. The model flags 120: 80 actual churners (TP), 40 non-churners (FP). Precision = 80 ÷ 120 = 67%; recall = 80 ÷ 100 = 80%; accuracy = (80 + 860) ÷ 1,000 = 94%.

  • ROC curve and AUC: trade-off between sensitivity and false positive rate across cut-offs (pROC package).
  • Imbalanced data: accuracy misleads — use precision, recall, F1, AUC; rebalance with oversampling (SMOTE) or class weights.
5

Topic 5

Benford's Law for fraud detection

  • Benford's Law: in many naturally occurring data sets, the leading digit d appears with probability log10(1 + 1/d) — 1 appears about 30.1% of the time, 9 about 4.6%.
Key formulasBenford's expected frequencies
  • Digit 1

    30.1%

  • Digit 2

    17.6%

  • Digit 3

    12.5%

  • Digit 5

    7.9%

  • Digit 9

    4.6%

  • Use: auditors compare actual first-digit frequencies of invoices, expenses or claims with Benford's distribution (chi-square test); large deviations flag possible manipulation.
  • Limits: not applicable to assigned numbers, data with fixed ranges or thresholds.

Example

An expense audit finds an unusual excess of claims starting with 4 and 9 just below approval limits of ₹5,000 and ₹10,000 — a red flag for splitting or fabricating claims.

6

Topic 6

K-means clustering for customer segmentation

  • Variables: balances, transactions, product holdings, channel use, age, income, credit usage.
  • Process: standardise variables, choose k (elbow, silhouette), run K-means, profile clusters, design strategies for each.

Example

A bank finds four clusters — young digital spenders, affluent investors, salaried savers and small-business owners — and tailors products and channels to each.

Key terms

Beta
Sensitivity of a stock's return to market return
Alpha
Return above that explained by risk factors
Probability of default
Likelihood a borrower defaults
Benford's Law
Distribution of leading digits in natural data
Customer segmentation
Grouping customers with similar characteristics

Quick revision

  • Market model; Fama–French; beta and alpha.
  • Forecasting financial variables; out-of-sample testing.
  • Logistic regression for PD; scorecards.
  • Confusion matrix, precision, recall; ROC–AUC.
  • Benford's Law frequencies and uses; K-means segmentation.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.What does beta measure?
  2. Q2.State the Fama–French three factors.
  3. Q3.Why is logistic regression used in credit scoring?
  4. Q4.What does AUC measure?
  5. Q5.State Benford's probability for the digit 1.
  6. Q6.What is K-means clustering?

Long-answer questions

  1. Q1.Explain regression and factor models in finance.
  2. Q2.Explain credit scoring using logistic regression and its evaluation.
  3. Q3.Explain Benford's Law and its use in fraud detection.
  4. Q4.Explain customer segmentation using K-means clustering.

Stuck on this unit?

Message SBS on WhatsApp for help with Financial Analytics, or to ask about studying MBA at Synetic.

WhatsApp us