Unit 4: Prediction and clustering
Data Mining for Business Decisions notes · PTU syllabus (MBA 941-18)
On this page
Unit summary
Prediction estimates numbers and clustering finds natural groups. This unit covers linear and multiple regression for prediction, types of data in cluster analysis — interval-scaled, binary, nominal, ordinal and ratio-scaled — partitioning methods (K-means and K-medoids), hierarchical agglomerative methods and DBSCAN.
After this unit you can
- Apply linear and multiple regression for prediction
- Measure similarity for different types of data in cluster analysis
- Apply K-means and K-medoids partitioning methods
- Explain hierarchical agglomerative clustering and DBSCAN
PTU syllabus topics
- Linear and multiple regression prediction
- cluster analysis data types (interval-scaled, binary, nominal, ordinal, ratio-scaled)
- partitioning methods (K-Means, K-Medoids)
- hierarchical agglomerative methods
- DBSCAN
K-means
Assign to nearest mean, update means
Fast; needs k
K-medoids
Uses actual points as centres
Robust to outliers
Hierarchical agglomerative
Merge closest clusters step by step
Gives a dendrogram
DBSCAN
Grows dense regions
Finds odd shapes and noise
Topic 1
Linear and multiple regression for prediction
Simple linear
Y = a + bX
Slope
b = [nΣXY − ΣXΣY] ÷ [nΣX² − (ΣX)²]
Intercept
a = Ȳ − bX̄
Multiple
Y = b0 + b1X1 + b2X2 + … + bkXk
Example
A model of monthly sales gives Sales = 20 + 3.5 × Ad spend (₹ lakh). At ₹8 lakh ad spend, predicted sales = 20 + 28 = ₹48 lakh.
- Evaluation: R², adjusted R², RMSE, MAPE on test data; check residuals and multicollinearity.
- Non-linear prediction: polynomial regression, log transformations, regression trees.
Topic 2
Types of data in cluster analysis
Interval-scaled
Standardise, then Euclidean or Manhattan distance
Binary
Simple matching coefficient (symmetric), Jaccard coefficient (asymmetric)
Nominal
Proportion of mismatches, or convert to binary
Ordinal
Replace by rank, scale to 0–1, treat as interval
Ratio-scaled
Log transformation, then interval methods
Mixed types
Weighted combination (Gower's distance)
Euclidean
√Σ(xi − yi)²
Manhattan
Σ abs(xi − yi)
Jaccard distance (binary)
(r + s) ÷ (q + r + s), ignoring negative matches
Topic 3
Partitioning methods: K-means and K-medoids
- 1Choose k
- 2Initialise k centroids
Randomly or by K-means++
- 3Assign each point to the nearest centroid
- 4Recompute centroids as cluster means
- 5Repeat until assignments stop changing
- Choosing k: elbow method (within-cluster sum of squares), silhouette score, business judgement.
- Limitations of K-means: sensitive to outliers and initial centroids, assumes roughly spherical clusters, numeric data only.
- K-medoids (PAM — partitioning around medoids): uses actual data points (medoids) as centres and minimises total dissimilarity — more robust to outliers; CLARA applies it to samples of large data.
Example
Customers plotted on annual spend and visit frequency form three K-means clusters: occasional low spenders, frequent mid spenders and a small group of high-value loyalists.
Topic 4
Hierarchical agglomerative methods
- Agglomerative (bottom-up): start with each point as a cluster and repeatedly merge the two closest clusters; divisive (top-down) splits from one cluster.
Single linkage
Distance between closest points — chaining effect
Complete linkage
Distance between farthest points — compact clusters
Average linkage
Average of all pairwise distances
Ward's method
Merge that least increases within-cluster variance
- Dendrogram: tree diagram of merges; cut it at a chosen height to get clusters.
- Pros and cons: no need to fix k in advance; but computationally expensive for large data and merges cannot be undone.
Topic 5
DBSCAN
- Density-based spatial clustering of applications with noise: clusters are dense regions separated by sparse regions.
- Parameters: eps (neighbourhood radius) and MinPts (minimum points to form a dense region).
- Point types: core points (at least MinPts within eps), border points (within eps of a core point), noise (neither).
- Strengths: finds clusters of arbitrary shape, identifies outliers, no need to specify k; weakness: struggles with varying densities and high dimensions.
K-means
Fast, simple, scalable
Needs k; sensitive to outliers; spherical clusters
K-medoids
Robust to outliers
Slower on large data
Hierarchical
Dendrogram; no k needed upfront
Slow; irreversible merges
DBSCAN
Arbitrary shapes; detects noise
Sensitive to eps and MinPts; varying density
Key terms
- Centroid
- Mean point of a cluster
- Medoid
- Most central actual data point of a cluster
- Dendrogram
- Tree diagram of hierarchical clustering
- Jaccard coefficient
- Similarity measure for asymmetric binary data
- Core point
- Point with at least MinPts neighbours within eps
Quick revision
- Regression for prediction; evaluation by R², RMSE, MAPE.
- Data types and similarity: interval, binary (matching, Jaccard), nominal, ordinal, ratio, mixed.
- K-means steps, choosing k; K-medoids (PAM), CLARA.
- Agglomerative clustering; single, complete, average, Ward linkage; dendrogram.
- DBSCAN: eps, MinPts, core, border, noise; comparison of methods.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.State the multiple regression equation.
- Q2.Distinguish Euclidean and Manhattan distance.
- Q3.When is the Jaccard coefficient used?
- Q4.State the steps of the K-means algorithm.
- Q5.Distinguish K-means and K-medoids.
- Q6.What are the parameters of DBSCAN?
Long-answer questions
- Q1.Explain linear and multiple regression for prediction with an example.
- Q2.Explain the types of data in cluster analysis and how similarity is measured.
- Q3.Explain K-means and K-medoids clustering.
- Q4.Explain hierarchical agglomerative clustering and DBSCAN.
Stuck on this unit?
Message SBS on WhatsApp for help with Data Mining for Business Decisions, or to ask about studying MBA at Synetic.
