Unit 3 of 4 · M.Sc IT Sem 4

Unit 3: Ensemble methods and clustering

Machine Learning notes · PTU syllabus (PGCA1945)

3 min read5 topics10 exam questions
On this page
  1. Unit summary
  2. Decision trees
  3. Ensemble methods
  4. Clustering
  5. Support vector machines
  6. Principal component analysis
  7. Key terms
  8. Quick revision
  9. Important questions

Unit summary

Combining models and finding structure in unlabelled data are powerful ideas. This unit covers decision trees, bagging and random forests, boosting, k-means and hierarchical clustering, support vector machines and principal component analysis.

After this unit you can

  • Build decision trees with information gain or Gini
  • Explain bagging, random forests and boosting
  • Apply k-means and hierarchical clustering
  • Explain SVMs and PCA

PTU syllabus topics

  • Bagging
  • decision trees and random forests
  • boosting
  • clustering (k-means, hierarchical agglomeration)
  • Support Vector Machines
  • Principal Component Analysis
ComparisonBagging vs boosting
Bagging
Boosting

Training

Models in parallel on bootstrap samples

Models in sequence, fixing earlier errors

Reduces

Variance

Bias

Example

Random forest

AdaBoost, gradient boosting

Overfitting risk

Lower

Higher if over-trained

1

Topic 1

Decision trees

Key formulasSplit measures
  • Entropy = −Σ p log2 p

    Impurity (ID3, C4.5)

  • Information gain = Entropy(parent) − weighted Entropy(children)

    Choose highest

  • Gini = 1 − Σ p²

    Impurity (CART)

Example

10 samples, 5 yes and 5 no → entropy 1. A split gives (4 yes, 1 no) and (1 yes, 4 no), each with entropy 0.72 → gain = 1 − 0.72 = 0.28.

  • Pruning (max depth, min samples) controls overfitting.
2

Topic 2

Ensemble methods

ComparisonEnsembles
Bagging and random forest
Boosting

Training

Models trained in parallel on bootstrap samples

Models trained in sequence, each fixing previous errors

Combination

Vote or average

Weighted vote

Reduces

Variance

Bias

Examples

Random forest (also random feature subsets per split)

AdaBoost, Gradient Boosting, XGBoost

3

Topic 3

Clustering

Processk-means
  1. 1Choose k
  2. 2Initialise centroids (k-means++)
  3. 3Assign points to the nearest centroid
  4. 4Update centroids as means
  5. 5Repeat until stable
  • Choosing k: elbow method (within-cluster sum of squares) or silhouette score.
  • Hierarchical agglomeration: start with single points, repeatedly merge the closest clusters (single, complete, average or Ward linkage); cut the dendrogram.
4

Topic 4

Support vector machines

Key formulasSVM
  • Hyperplane wᵀx + b = 0

    Separator

  • Margin = 2 / norm(w)

    Maximised

  • Soft margin: minimise ½ norm(w)² + C Σ ξi

    C trades margin vs errors

  • Kernels: linear, polynomial, RBF e^(−γ dist²)

    Non-linear boundaries

5

Topic 5

Principal component analysis

ProcessPCA
  1. 1

    Standardise the data

  2. 2

    Compute the covariance matrix

  3. 3

    Find eigenvectors and eigenvalues

  4. 4

    Sort by eigenvalue (variance explained)

  5. 5

    Keep top k components

  6. 6

    Project data onto them

  • Uses: dimensionality reduction, visualisation in 2-D, noise removal, faster training. Choose k to retain, say, 95% of variance.

Key terms

Information gain
Reduction in entropy from a split
Random forest
Bagged trees with random feature subsets
Boosting
Sequential ensemble focusing on errors
Support vector
Training point on the margin
Principal component
Direction of maximum variance

Quick revision

  • Entropy, information gain, Gini; pruning.
  • Bagging reduces variance; boosting reduces bias.
  • k-means, elbow, silhouette; hierarchical linkage.
  • SVM margin, soft margin, kernels; PCA steps.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.Define entropy.
  2. Q2.What is bagging?
  3. Q3.How does boosting differ from bagging?
  4. Q4.What is the elbow method?
  5. Q5.What is the kernel trick?
  6. Q6.What is PCA used for?

Long-answer questions

  1. Q1.Explain decision tree construction with an example.
  2. Q2.Explain bagging, random forests and boosting.
  3. Q3.Explain SVMs.
  4. Q4.Explain principal component analysis.

Stuck on this unit?

Message SBS on WhatsApp for help with Machine Learning, or to ask about studying M.Sc IT at Synetic.

WhatsApp us