Unit 1 of 4 · M.Sc IT Sem 4

Unit 1: ML fundamentals and regression

Machine Learning notes · PTU syllabus (PGCA1945)

3 min read9 topics9 exam questions
On this page
  1. Unit summary
  2. What is machine learning?
  3. Types of machine learning
  4. Problems, data and tools
  5. Performance evaluation measures
  6. Error metrics for regression
  7. Data visualisation
  8. Linear regression
  9. Overfitting and data splits
  10. Training, validation and test data
  11. Key terms
  12. Quick revision
  13. Important questions

Unit summary

Machine learning lets computers learn patterns from data instead of being explicitly programmed. This unit covers the ML problem setting, types of learning, evaluation measures, data visualisation, and linear regression solved by gradient descent and the normal equation, with features, overfitting and data splits.

After this unit you can

  • Explain machine learning, its problems, data and tools
  • Compare types of learning
  • Evaluate models with accuracy, precision, recall, F-measure and error metrics
  • Fit linear regression by gradient descent and the normal equation

PTU syllabus topics

  • What is machine learning
  • problems/data/tools
  • types of learning
  • performance evaluation measures (accuracy, precision, recall, F-measure)
  • error metrics
  • data visualization
  • linear regression
  • gradient descent
  • closed-form/normal equations
  • features
  • overfitting
  • training/validation/test data
Key formulasLinear regression
  • Hypothesis

    ŷ = w₀ + w₁x₁ + … + wₙxₙ

  • Cost (MSE)

    J = (1/2m) Σ (ŷ − y)²

  • Gradient descent

    w ← w − α ∂J/∂w

  • Normal equation

    w = (XᵀX)⁻¹ Xᵀ y

1

Topic 1

What is machine learning?

Tom Mitchell's definition: a program learns from experience E with respect to task T and performance measure P if its performance on T, measured by P, improves with E.

Example

Spam filter — T: classify emails; E: emails labelled spam or not; P: percentage classified correctly.

History in brief: perceptron (1958), decision trees and backpropagation (1980s), SVMs and random forests (1990s–2000s), deep learning breakthroughs (2012 onward), and today's large language models. Applications: recommendations (Netflix, Amazon), fraud detection, medical diagnosis, speech recognition, self-driving cars and price prediction.

2

Topic 2

Types of machine learning

ComparisonTypes of learning
Data used
Example

Supervised

Labelled data (input + correct output)

Predict house prices, classify emails

Unsupervised

Unlabelled data

Customer segmentation, clustering

Semi-supervised

A little labelled + lots of unlabelled

Photo tagging with few labels

Reinforcement

Rewards and penalties from an environment

Game-playing agents, robotics

Regression predicts a continuous number (price, temperature); classification predicts a category (spam/not spam, pass/fail).

3

Topic 3

Problems, data and tools

Key termsML ingredients
Problem
Regression, classification, clustering, ranking, forecasting
Data
Features (inputs) and labels (targets); tabular, text, images
Tools
Python (NumPy, pandas, scikit-learn, TensorFlow, PyTorch), R, MATLAB, Weka
4

Topic 4

Performance evaluation measures

Predicted positivePredicted negative
Actual positiveTrue Positive (TP)False Negative (FN)
Actual negativeFalse Positive (FP)True Negative (TN)
Key formulasClassification metrics
  • Accuracy

    (TP + TN) / total

  • Precision

    TP / (TP + FP)

    Of predicted positives, how many are right

  • Recall (sensitivity)

    TP / (TP + FN)

    Of actual positives, how many were found

  • F1 score

    2 × Precision × Recall / (Precision + Recall)

  • AUC

    Area under the ROC curve; 1.0 is perfect, 0.5 is random

Example

TP = 40, FP = 10, FN = 20, TN = 30. Accuracy = 70/100 = 70%; precision = 40/50 = 80%; recall = 40/60 ≈ 67%.

Exam tip

Accuracy misleads on imbalanced data — a model that always says "no fraud" is 99% accurate if fraud is 1%. Use precision, recall and F1.

5

Topic 5

Error metrics for regression

Key formulasError metrics
  • MAE = (1/n) Σ abs(y − ŷ)

    Mean absolute error

  • MSE = (1/n) Σ (y − ŷ)²

    Mean squared error

  • RMSE = √MSE

    Same units as y

  • R² = 1 − SS_res / SS_tot

    Proportion of variance explained

6

Topic 6

Data visualisation

Key termsPlots
Histogram
Distribution of one feature
Scatter plot
Relationship between two features
Box plot
Spread and outliers
Heat map
Correlation matrix
Pair plot
All pairwise scatters
7

Topic 7

Linear regression

Key formulasLinear regression
  • ŷ = θ0 + θ1x1 + … + θnxn

    Hypothesis

  • J(θ) = (1/2m) Σ (ŷ − y)²

    Cost function

  • θj = θj − α × (1/m) Σ (ŷ − y) xj

    Gradient descent update

  • θ = (XᵀX)⁻¹ Xᵀy

    Normal (closed-form) equation

ComparisonSolving for θ
Gradient descent
Normal equation

Method

Iterative; choose learning rate α

Direct formula

Large n features

Scales well

Slow — inverting XᵀX is O(n³)

Feature scaling

Needed

Not needed

Example

Data (x, y) = (1, 2), (2, 4), (3, 6): the normal equation gives θ0 = 0, θ1 = 2, so ŷ = 2x and J = 0.

  • Features: feature engineering (polynomial terms, interactions), scaling (standardisation, min–max), encoding categories (one-hot).
8

Topic 8

Overfitting and data splits

ComparisonFit
Underfitting
Overfitting

Cause

Model too simple

Model too complex for the data

Symptom

High training and test error

Low training error, high test error

Remedy

More features, complex model

More data, regularisation (L1 lasso, L2 ridge), simpler model, early stopping

9

Topic 9

Training, validation and test data

ProcessThe ML workflow
  1. 1

    Collect and clean data

  2. 2

    Split data

    e.g. 70% train, 15% validation, 15% test

  3. 3

    Train the model

    On the training set

  4. 4

    Tune hyperparameters

    Using the validation set

  5. 5

    Evaluate once

    On the unseen test set

  6. 6

    Deploy and monitor

  • Overfitting: the model memorises training data and performs poorly on new data (high variance).
  • Underfitting: the model is too simple to capture the pattern (high bias).
  • Cross-validation (k-fold) gives a more reliable estimate by rotating the validation fold.

Key terms

Feature
Input variable
Cost function
Measures model error to be minimised
Learning rate
Step size in gradient descent
Overfitting
Fitting noise so the model fails on new data
Regularisation
Penalty on large weights to reduce overfitting

Quick revision

  • Supervised, unsupervised, reinforcement learning.
  • Accuracy, precision, recall, F1; MAE, MSE, RMSE, R².
  • Hypothesis, cost, gradient descent, normal equation.
  • Over- and underfitting; regularisation; train, validation, test; cross-validation.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.Define machine learning.
  2. Q2.Distinguish supervised and unsupervised learning.
  3. Q3.Define precision and recall.
  4. Q4.What is gradient descent?
  5. Q5.State the normal equation.
  6. Q6.What is overfitting?

Long-answer questions

  1. Q1.Explain the types of machine learning with examples.
  2. Q2.Explain performance evaluation measures for classification and regression.
  3. Q3.Explain linear regression with gradient descent and the normal equation.

Stuck on this unit?

Message SBS on WhatsApp for help with Machine Learning, or to ask about studying M.Sc IT at Synetic.

WhatsApp us