Unit 4 of 4 · MBA Sem 3

Unit 4: Prediction and clustering

Data Mining for Business Decisions notes · PTU syllabus (MBA 941-18)

3 min read5 topics10 exam questions
On this page
  1. Unit summary
  2. Linear and multiple regression for prediction
  3. Types of data in cluster analysis
  4. Partitioning methods: K-means and K-medoids
  5. Hierarchical agglomerative methods
  6. DBSCAN
  7. Key terms
  8. Quick revision
  9. Important questions

Unit summary

Prediction estimates numbers and clustering finds natural groups. This unit covers linear and multiple regression for prediction, types of data in cluster analysis — interval-scaled, binary, nominal, ordinal and ratio-scaled — partitioning methods (K-means and K-medoids), hierarchical agglomerative methods and DBSCAN.

After this unit you can

  • Apply linear and multiple regression for prediction
  • Measure similarity for different types of data in cluster analysis
  • Apply K-means and K-medoids partitioning methods
  • Explain hierarchical agglomerative clustering and DBSCAN

PTU syllabus topics

  • Linear and multiple regression prediction
  • cluster analysis data types (interval-scaled, binary, nominal, ordinal, ratio-scaled)
  • partitioning methods (K-Means, K-Medoids)
  • hierarchical agglomerative methods
  • DBSCAN
ComparisonClustering methods
How
Note

K-means

Assign to nearest mean, update means

Fast; needs k

K-medoids

Uses actual points as centres

Robust to outliers

Hierarchical agglomerative

Merge closest clusters step by step

Gives a dendrogram

DBSCAN

Grows dense regions

Finds odd shapes and noise

1

Topic 1

Linear and multiple regression for prediction

Key formulasRegression models
  • Simple linear

    Y = a + bX

  • Slope

    b = [nΣXY − ΣXΣY] ÷ [nΣX² − (ΣX)²]

  • Intercept

    a = Ȳ − bX̄

  • Multiple

    Y = b0 + b1X1 + b2X2 + … + bkXk

Example

A model of monthly sales gives Sales = 20 + 3.5 × Ad spend (₹ lakh). At ₹8 lakh ad spend, predicted sales = 20 + 28 = ₹48 lakh.

  • Evaluation: R², adjusted R², RMSE, MAPE on test data; check residuals and multicollinearity.
  • Non-linear prediction: polynomial regression, log transformations, regression trees.
2

Topic 2

Types of data in cluster analysis

ClassificationSimilarity measures by data type
Data types
  • Interval-scaled

    Standardise, then Euclidean or Manhattan distance

  • Binary

    Simple matching coefficient (symmetric), Jaccard coefficient (asymmetric)

  • Nominal

    Proportion of mismatches, or convert to binary

  • Ordinal

    Replace by rank, scale to 0–1, treat as interval

  • Ratio-scaled

    Log transformation, then interval methods

  • Mixed types

    Weighted combination (Gower's distance)

Key formulasDistance measures
  • Euclidean

    √Σ(xi − yi)²

  • Manhattan

    Σ abs(xi − yi)

  • Jaccard distance (binary)

    (r + s) ÷ (q + r + s), ignoring negative matches

3

Topic 3

Partitioning methods: K-means and K-medoids

ProcessK-means algorithm
  1. 1Choose k
  2. 2Initialise k centroids

    Randomly or by K-means++

  3. 3Assign each point to the nearest centroid
  4. 4Recompute centroids as cluster means
  5. 5Repeat until assignments stop changing
  • Choosing k: elbow method (within-cluster sum of squares), silhouette score, business judgement.
  • Limitations of K-means: sensitive to outliers and initial centroids, assumes roughly spherical clusters, numeric data only.
  • K-medoids (PAM — partitioning around medoids): uses actual data points (medoids) as centres and minimises total dissimilarity — more robust to outliers; CLARA applies it to samples of large data.

Example

Customers plotted on annual spend and visit frequency form three K-means clusters: occasional low spenders, frequent mid spenders and a small group of high-value loyalists.

4

Topic 4

Hierarchical agglomerative methods

  • Agglomerative (bottom-up): start with each point as a cluster and repeatedly merge the two closest clusters; divisive (top-down) splits from one cluster.
ClassificationLinkage methods
Linkage
  • Single linkage

    Distance between closest points — chaining effect

  • Complete linkage

    Distance between farthest points — compact clusters

  • Average linkage

    Average of all pairwise distances

  • Ward's method

    Merge that least increases within-cluster variance

  • Dendrogram: tree diagram of merges; cut it at a chosen height to get clusters.
  • Pros and cons: no need to fix k in advance; but computationally expensive for large data and merges cannot be undone.
5

Topic 5

DBSCAN

  • Density-based spatial clustering of applications with noise: clusters are dense regions separated by sparse regions.
  • Parameters: eps (neighbourhood radius) and MinPts (minimum points to form a dense region).
  • Point types: core points (at least MinPts within eps), border points (within eps of a core point), noise (neither).
  • Strengths: finds clusters of arbitrary shape, identifies outliers, no need to specify k; weakness: struggles with varying densities and high dimensions.
ComparisonClustering methods compared
Strength
Limitation

K-means

Fast, simple, scalable

Needs k; sensitive to outliers; spherical clusters

K-medoids

Robust to outliers

Slower on large data

Hierarchical

Dendrogram; no k needed upfront

Slow; irreversible merges

DBSCAN

Arbitrary shapes; detects noise

Sensitive to eps and MinPts; varying density

Key terms

Centroid
Mean point of a cluster
Medoid
Most central actual data point of a cluster
Dendrogram
Tree diagram of hierarchical clustering
Jaccard coefficient
Similarity measure for asymmetric binary data
Core point
Point with at least MinPts neighbours within eps

Quick revision

  • Regression for prediction; evaluation by R², RMSE, MAPE.
  • Data types and similarity: interval, binary (matching, Jaccard), nominal, ordinal, ratio, mixed.
  • K-means steps, choosing k; K-medoids (PAM), CLARA.
  • Agglomerative clustering; single, complete, average, Ward linkage; dendrogram.
  • DBSCAN: eps, MinPts, core, border, noise; comparison of methods.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.State the multiple regression equation.
  2. Q2.Distinguish Euclidean and Manhattan distance.
  3. Q3.When is the Jaccard coefficient used?
  4. Q4.State the steps of the K-means algorithm.
  5. Q5.Distinguish K-means and K-medoids.
  6. Q6.What are the parameters of DBSCAN?

Long-answer questions

  1. Q1.Explain linear and multiple regression for prediction with an example.
  2. Q2.Explain the types of data in cluster analysis and how similarity is measured.
  3. Q3.Explain K-means and K-medoids clustering.
  4. Q4.Explain hierarchical agglomerative clustering and DBSCAN.

Stuck on this unit?

Message SBS on WhatsApp for help with Data Mining for Business Decisions, or to ask about studying MBA at Synetic.

WhatsApp us