Unit 3: Ensemble methods and clustering
Data Sciences Using R notes · PTU syllabus (MBA 962-18)
On this page
Unit summary
Combining many models often beats any single model, and clustering reveals hidden groups. This unit covers ensemble methods — bagging, random forests and boosting — and clustering methods — K-means, K-medoids, agglomerative and hierarchical clustering, X-means and DBSCAN.
After this unit you can
- Explain bagging and random forests
- Explain boosting
- Apply K-means, K-medoids and X-means clustering
- Apply hierarchical clustering and DBSCAN
PTU syllabus topics
- Bagging
- random forests
- boosting
- K-means clustering
- K-medoids
- agglomerative and hierarchical clustering
- X-means
- DBSCAN
Random forest
Many trees on bootstrap samples
randomForest()
Boosting
Trees in sequence fixing errors
gbm, xgboost
K-means
Assign to nearest centre
kmeans()
Hierarchical
Merge clusters into a dendrogram
hclust()
DBSCAN
Density-based clusters
dbscan package
Topic 1
Ensemble methods
Ensemble learning combines several models to improve accuracy and stability.
Training
Models trained in parallel on bootstrap samples
Models trained sequentially, each correcting previous errors
Combination
Average or majority vote
Weighted vote
Main effect
Reduces variance (overfitting)
Reduces bias
Examples
Bagged trees, random forest
AdaBoost, gradient boosting, XGBoost
Topic 2
Bagging and random forests
- Bagging (bootstrap aggregating, Breiman 1996): draw many bootstrap samples, fit a model (usually a tree) to each, and combine predictions.
- Random forest: bagging plus a random subset of variables at each split, making trees less correlated.
- In R: randomForest(default ~ ., data = train, ntree = 500, mtry = 3); importance(model) and varImpPlot(model) show variable importance.
- Out-of-bag (OOB) error: estimate of test error using observations left out of each bootstrap sample.
Topic 3
Boosting
- AdaBoost: increases the weights of misclassified cases so later models focus on them.
- Gradient boosting: each new tree fits the residual errors of the current ensemble; XGBoost adds regularisation and speed — popular in competitions and credit scoring.
- In R: gbm, xgboost packages; tune learning rate, number of trees and depth to avoid overfitting.
Topic 4
K-means and K-medoids
- K-means: kmeans(scale(df), centers = 4, nstart = 25) — assigns points to the nearest of k centroids and updates centroids until stable.
- Choosing k: elbow method (within-cluster sum of squares), silhouette width (cluster::silhouette), gap statistic.
- K-medoids (PAM): cluster::pam(df, k = 4) — uses actual observations as centres; more robust to outliers.
Topic 5
X-means
- X-means (Pelleg and Moore, 2000): an extension of K-means that chooses the number of clusters automatically by repeatedly splitting clusters and keeping splits that improve the Bayesian Information Criterion (BIC).
- Useful when k is unknown and data are large.
Topic 6
Agglomerative and hierarchical clustering
- Agglomerative: start with each observation as a cluster and merge the closest pairs; divisive splits from one cluster.
- In R: d <- dist(scale(df)); hc <- hclust(d, method = "ward.D2"); plot(hc); cutree(hc, k = 4).
- Linkage: single, complete, average, Ward; the dendrogram shows merge order and heights.
Topic 7
DBSCAN
- Density-based clustering: groups points in dense regions; points in sparse regions are noise.
- Parameters: eps (neighbourhood radius) and minPts; choose eps from the k-nearest-neighbour distance plot (dbscan::kNNdistplot).
- In R: dbscan::dbscan(df, eps = 0.5, minPts = 5).
- Strengths: arbitrary-shaped clusters and outlier detection without fixing k.
K-means
Fixed in advance
Large data, roughly spherical clusters
K-medoids
Fixed in advance
Data with outliers
X-means
Chosen by BIC
Unknown k, large data
Hierarchical
Chosen from dendrogram
Small data, nested structure
DBSCAN
Found from density
Irregular shapes, noise
Key terms
- Ensemble
- Combination of multiple models
- Bootstrap sample
- Sample drawn with replacement
- Random forest
- Bagged trees with random variable selection at splits
- Gradient boosting
- Sequential trees fitting residual errors
- X-means
- K-means variant choosing k by BIC
Quick revision
- Bagging (variance) vs boosting (bias).
- Random forest: ntree, mtry, OOB error, variable importance.
- AdaBoost, gradient boosting, XGBoost.
- K-means (elbow, silhouette), K-medoids (PAM), X-means (BIC).
- Hierarchical (hclust, linkage, dendrogram); DBSCAN (eps, minPts).
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.What is ensemble learning?
- Q2.Distinguish bagging and boosting.
- Q3.What is out-of-bag error?
- Q4.How does X-means choose the number of clusters?
- Q5.Write the R command for hierarchical clustering.
- Q6.What are the parameters of DBSCAN?
Long-answer questions
- Q1.Explain bagging and random forests.
- Q2.Explain boosting methods and their advantages.
- Q3.Explain K-means, K-medoids and X-means clustering.
- Q4.Explain hierarchical clustering and DBSCAN with R implementation.
Stuck on this unit?
Message SBS on WhatsApp for help with Data Sciences Using R, or to ask about studying MBA at Synetic.
