Unit 3 of 4 · MBA Sem 3

Unit 3: Ensemble methods and clustering

Data Sciences Using R notes · PTU syllabus (MBA 962-18)

3 min read7 topics10 exam questions
On this page
  1. Unit summary
  2. Ensemble methods
  3. Bagging and random forests
  4. Boosting
  5. K-means and K-medoids
  6. X-means
  7. Agglomerative and hierarchical clustering
  8. DBSCAN
  9. Key terms
  10. Quick revision
  11. Important questions

Unit summary

Combining many models often beats any single model, and clustering reveals hidden groups. This unit covers ensemble methods — bagging, random forests and boosting — and clustering methods — K-means, K-medoids, agglomerative and hierarchical clustering, X-means and DBSCAN.

After this unit you can

  • Explain bagging and random forests
  • Explain boosting
  • Apply K-means, K-medoids and X-means clustering
  • Apply hierarchical clustering and DBSCAN

PTU syllabus topics

  • Bagging
  • random forests
  • boosting
  • K-means clustering
  • K-medoids
  • agglomerative and hierarchical clustering
  • X-means
  • DBSCAN
ComparisonEnsemble and clustering methods
Idea
R function or package

Random forest

Many trees on bootstrap samples

randomForest()

Boosting

Trees in sequence fixing errors

gbm, xgboost

K-means

Assign to nearest centre

kmeans()

Hierarchical

Merge clusters into a dendrogram

hclust()

DBSCAN

Density-based clusters

dbscan package

1

Topic 1

Ensemble methods

Ensemble learning combines several models to improve accuracy and stability.

ComparisonBagging vs boosting
Bagging
Boosting

Training

Models trained in parallel on bootstrap samples

Models trained sequentially, each correcting previous errors

Combination

Average or majority vote

Weighted vote

Main effect

Reduces variance (overfitting)

Reduces bias

Examples

Bagged trees, random forest

AdaBoost, gradient boosting, XGBoost

2

Topic 2

Bagging and random forests

  • Bagging (bootstrap aggregating, Breiman 1996): draw many bootstrap samples, fit a model (usually a tree) to each, and combine predictions.
  • Random forest: bagging plus a random subset of variables at each split, making trees less correlated.
  • In R: randomForest(default ~ ., data = train, ntree = 500, mtry = 3); importance(model) and varImpPlot(model) show variable importance.
  • Out-of-bag (OOB) error: estimate of test error using observations left out of each bootstrap sample.
3

Topic 3

Boosting

  • AdaBoost: increases the weights of misclassified cases so later models focus on them.
  • Gradient boosting: each new tree fits the residual errors of the current ensemble; XGBoost adds regularisation and speed — popular in competitions and credit scoring.
  • In R: gbm, xgboost packages; tune learning rate, number of trees and depth to avoid overfitting.
4

Topic 4

K-means and K-medoids

  • K-means: kmeans(scale(df), centers = 4, nstart = 25) — assigns points to the nearest of k centroids and updates centroids until stable.
  • Choosing k: elbow method (within-cluster sum of squares), silhouette width (cluster::silhouette), gap statistic.
  • K-medoids (PAM): cluster::pam(df, k = 4) — uses actual observations as centres; more robust to outliers.
5

Topic 5

X-means

  • X-means (Pelleg and Moore, 2000): an extension of K-means that chooses the number of clusters automatically by repeatedly splitting clusters and keeping splits that improve the Bayesian Information Criterion (BIC).
  • Useful when k is unknown and data are large.
6

Topic 6

Agglomerative and hierarchical clustering

  • Agglomerative: start with each observation as a cluster and merge the closest pairs; divisive splits from one cluster.
  • In R: d <- dist(scale(df)); hc <- hclust(d, method = "ward.D2"); plot(hc); cutree(hc, k = 4).
  • Linkage: single, complete, average, Ward; the dendrogram shows merge order and heights.
7

Topic 7

DBSCAN

  • Density-based clustering: groups points in dense regions; points in sparse regions are noise.
  • Parameters: eps (neighbourhood radius) and minPts; choose eps from the k-nearest-neighbour distance plot (dbscan::kNNdistplot).
  • In R: dbscan::dbscan(df, eps = 0.5, minPts = 5).
  • Strengths: arbitrary-shaped clusters and outlier detection without fixing k.
ComparisonClustering methods
Number of clusters
Best for

K-means

Fixed in advance

Large data, roughly spherical clusters

K-medoids

Fixed in advance

Data with outliers

X-means

Chosen by BIC

Unknown k, large data

Hierarchical

Chosen from dendrogram

Small data, nested structure

DBSCAN

Found from density

Irregular shapes, noise

Key terms

Ensemble
Combination of multiple models
Bootstrap sample
Sample drawn with replacement
Random forest
Bagged trees with random variable selection at splits
Gradient boosting
Sequential trees fitting residual errors
X-means
K-means variant choosing k by BIC

Quick revision

  • Bagging (variance) vs boosting (bias).
  • Random forest: ntree, mtry, OOB error, variable importance.
  • AdaBoost, gradient boosting, XGBoost.
  • K-means (elbow, silhouette), K-medoids (PAM), X-means (BIC).
  • Hierarchical (hclust, linkage, dendrogram); DBSCAN (eps, minPts).

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.What is ensemble learning?
  2. Q2.Distinguish bagging and boosting.
  3. Q3.What is out-of-bag error?
  4. Q4.How does X-means choose the number of clusters?
  5. Q5.Write the R command for hierarchical clustering.
  6. Q6.What are the parameters of DBSCAN?

Long-answer questions

  1. Q1.Explain bagging and random forests.
  2. Q2.Explain boosting methods and their advantages.
  3. Q3.Explain K-means, K-medoids and X-means clustering.
  4. Q4.Explain hierarchical clustering and DBSCAN with R implementation.

Stuck on this unit?

Message SBS on WhatsApp for help with Data Sciences Using R, or to ask about studying MBA at Synetic.

WhatsApp us