Unit 2 of 2 · BCA Sem 4

Unit 2: Supervised and unsupervised learning

Introduction to Machine Learning notes · PTU syllabus (UGDSE203)

3 min read5 topics10 exam questions
On this page
  1. Unit summary
  2. Regression
  3. Naive Bayes, k-NN and decision trees
  4. Neural networks, perceptron and SVM
  5. Unsupervised learning: clustering
  6. Ethics and real-world applications
  7. Key terms
  8. Quick revision
  9. Important questions

Unit summary

This unit surveys the core algorithms of supervised and unsupervised learning: linear, non-linear and logistic regression, Naive Bayes, k-nearest neighbours, decision trees, the perceptron and neural networks, support vector machines, and clustering with K-means, hierarchical clustering and DBSCAN, along with the ethics of ML.

After this unit you can

  • Explain regression algorithms and logistic regression
  • Explain Naive Bayes, k-NN and decision trees
  • Describe the perceptron, single-layer networks and SVMs
  • Apply K-means, hierarchical clustering and DBSCAN and evaluate clusters

PTU syllabus topics

  • Linear and non-linear regression
  • logistic regression
  • Naive Bayes
  • K-Nearest Neighbors
  • decision trees
  • introduction to artificial neural networks
  • perceptron learning algorithm
  • single-layer perceptron
  • introduction to SVM for linearly separable data
  • K-Means
  • hierarchical clustering
  • DBSCAN
  • clustering validation measures
  • ethical considerations and real-world applications of ML
ClassificationMachine learning algorithms
Machine learning
  • Regression

    Linear and non-linear regression

  • Classification

    Logistic regression, Naive Bayes, KNN, decision trees

  • Neural networks

    Perceptron, SVM basics

  • Clustering

    K-Means, hierarchical, DBSCAN

1

Topic 1

Regression

  • Linear regression fits a straight line y = mx + c (or y = b₀ + b₁x₁ + … for many features) by minimising the sum of squared errors.
  • Non-linear (polynomial) regression fits curves using powers of x.
  • Logistic regression is a classification algorithm: it passes a linear score through the sigmoid function σ(z) = 1/(1 + e⁻ᶻ) to give a probability between 0 and 1, and predicts class 1 if the probability ≥ 0.5.
2

Topic 2

Naive Bayes, k-NN and decision trees

ComparisonClassic classifiers
Idea
Strength

Naive Bayes

Bayes' theorem assuming independent features

Fast; great for text and spam

k-Nearest Neighbours

Vote of the k closest training points

Simple; no training phase

Decision tree

If-then splits using information gain or Gini

Easy to interpret

Example

k-NN with k = 3: a new point's three nearest neighbours are 2 "pass" and 1 "fail", so it is classified "pass".

3

Topic 3

Neural networks, perceptron and SVM

An artificial neuron computes a weighted sum of inputs plus a bias and applies an activation function. The perceptron learns with the rule w ← w + η (t − y) x.

  • A single-layer perceptron can learn linearly separable functions like AND and OR, but not XOR — which needs a multi-layer network.
  • A support vector machine (SVM) finds the separating line (hyperplane) with the maximum margin between classes; the closest points are support vectors.
4

Topic 4

Unsupervised learning: clustering

ComparisonClustering algorithms
How it works
Note

K-means

Assign points to the nearest of k centroids; recompute centroids; repeat

Fast; needs k; round clusters

Hierarchical

Merge (agglomerative) or split (divisive) clusters step by step

Gives a dendrogram; no k needed in advance

DBSCAN

Grows clusters from dense regions

Finds any shape; labels noise; no k needed

Cluster validation: the elbow method (plot within-cluster sum of squares against k), the silhouette score (from −1 to 1; higher is better) and the Davies-Bouldin index.

5

Topic 5

Ethics and real-world applications

  • Bias: models trained on biased data make unfair decisions (in hiring or lending).
  • Privacy: personal data must be collected with consent and protected.
  • Transparency: people affected should be able to understand decisions.
  • Accountability: humans remain responsible for ML outcomes.

Applications: credit scoring, disease prediction, demand forecasting, recommendation engines and customer segmentation.

Key terms

Linear regression
Fitting a straight line to predict a number
Sigmoid
Function mapping any number to a probability between 0 and 1
Perceptron
A single artificial neuron that learns linear boundaries
Support vectors
Points closest to the SVM decision boundary
Silhouette score
A measure of how well points fit their clusters

Quick revision

  • Logistic regression classifies using the sigmoid.
  • Naive Bayes assumes independent features; k-NN votes; trees split.
  • A single perceptron can't learn XOR.
  • SVM maximises the margin.
  • K-means needs k; DBSCAN finds noise and odd shapes.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.Why is logistic regression a classification algorithm?
  2. Q2.What is the "naive" assumption in Naive Bayes?
  3. Q3.Why can't a single-layer perceptron learn XOR?
  4. Q4.What is a support vector?
  5. Q5.Differentiate between K-means and DBSCAN.
  6. Q6.What is the elbow method?

Long-answer questions

  1. Q1.Explain linear and logistic regression with equations and examples.
  2. Q2.Explain Naive Bayes, k-NN and decision tree classifiers.
  3. Q3.Explain the perceptron learning algorithm and its limitation.
  4. Q4.Explain K-means, hierarchical clustering and DBSCAN, and how clusters are validated.

Stuck on this unit?

Message SBS on WhatsApp for help with Introduction to Machine Learning, or to ask about studying BCA at Synetic.

WhatsApp us