Unit 2 of 4 · M.Sc IT Sem 4

Unit 2: Classification

Machine Learning notes · PTU syllabus (PGCA1945)

3 min read5 topics9 exam questions
On this page
  1. Unit summary
  2. Classification and decision boundaries
  3. K-nearest neighbour
  4. Logistic regression
  5. Probability and Bayes optimal decisions
  6. Naive Bayes and Gaussian class-conditionals
  7. Key terms
  8. Quick revision
  9. Important questions

Unit summary

Classification predicts categories such as spam or not spam. This unit covers decision boundaries, k-nearest neighbour, logistic regression, probabilistic classification, Bayes optimal decisions and Naive Bayes with Gaussian class-conditional distributions.

After this unit you can

  • Explain classification problems and decision boundaries
  • Apply k-NN
  • Explain logistic regression and its cost function
  • Apply Bayes optimal decisions and Gaussian Naive Bayes

PTU syllabus topics

  • Classification problems and decision boundaries
  • K-Nearest Neighbor
  • logistic regression
  • probability and classification
  • Bayes optimal decisions
  • Naive Bayes and Gaussian class-conditional distribution
Key formulasClassification essentials
  • Sigmoid

    σ(z) = 1 / (1 + e^(−z))

  • Decision boundary

    Predict 1 if σ(z) ≥ 0.5

  • Naive Bayes

    P(c given x) ∝ P(c) Π P(xi given c)

  • k-NN

    Majority class among the k nearest points

1

Topic 1

Classification and decision boundaries

  • A classifier divides feature space into regions; the decision boundary separates them — linear (logistic regression, linear SVM) or non-linear (k-NN, trees, kernels).
2

Topic 2

K-nearest neighbour

Processk-NN
  1. 1Scale features
  2. 2Compute distance to all training points
  3. 3Take the k nearest
  4. 4Majority vote (weighted by 1/distance optionally)
  • Small k → noisy, complex boundary (overfit); large k → smooth boundary (underfit). Choose k by cross-validation.
3

Topic 3

Logistic regression

Key formulasLogistic regression
  • σ(z) = 1 / (1 + e^(−z)), z = θᵀx

    Sigmoid gives P(y = 1)

  • Predict 1 if σ(z) ≥ 0.5

    Threshold

  • J(θ) = −(1/m) Σ [y log ŷ + (1 − y) log(1 − ŷ)]

    Log loss (cross-entropy)

  • Softmax

    Multiclass extension

Example

z = −4 + 0.05 × marks; for marks 100, z = 1, P(pass) = 1/(1 + e^−1) ≈ 0.73 → predict pass.

4

Topic 4

Probability and Bayes optimal decisions

Key formulasBayes
  • P(C given x) = P(x given C) P(C) / P(x)

    Posterior

  • Bayes optimal: choose C maximising P(C given x)

    Minimises error probability

  • With costs: choose the class with least expected loss

    Risk-sensitive decisions

  • Generative models (Naive Bayes) learn P(x given C) and P(C); discriminative models (logistic regression) learn P(C given x) directly.
5

Topic 5

Naive Bayes and Gaussian class-conditionals

Key formulasGaussian Naive Bayes
  • P(xj given C) = (1 / √(2πσ²)) e^(−(xj − μ)² / 2σ²)

    Per-feature normal density

  • Score(C) = P(C) × Π P(xj given C)

    Pick the highest

Example

Heights of class "adult" μ = 165, σ = 10; "child" μ = 120, σ = 15; equal priors. For x = 150: z-values 1.5 and 2.0, so adult density is higher → predict adult.

pythonfrom sklearn.naive_bayes import GaussianNB
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
for m in (GaussianNB(), LogisticRegression(max_iter=500), KNeighborsClassifier(5)):
    print(type(m).__name__, m.fit(X_train, y_train).score(X_test, y_test))

Key terms

Decision boundary
Surface separating predicted classes
Sigmoid
Function mapping any value to 0–1
Log loss
Cost function of logistic regression
Bayes optimal classifier
Chooses the most probable class
Gaussian Naive Bayes
Naive Bayes with normal class-conditionals

Quick revision

  • Linear vs non-linear boundaries.
  • k-NN: scaling, choice of k.
  • Sigmoid, threshold, log loss, softmax.
  • Bayes rule, optimal decisions, generative vs discriminative; Gaussian NB.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.What is a decision boundary?
  2. Q2.How does k affect k-NN?
  3. Q3.Write the sigmoid function.
  4. Q4.Why not use squared error for logistic regression?
  5. Q5.What is the Bayes optimal decision?
  6. Q6.Distinguish generative and discriminative models.

Long-answer questions

  1. Q1.Explain k-NN classification.
  2. Q2.Explain logistic regression with its cost function.
  3. Q3.Explain Naive Bayes with Gaussian class-conditional distributions.

Stuck on this unit?

Message SBS on WhatsApp for help with Machine Learning, or to ask about studying M.Sc IT at Synetic.

WhatsApp us