Unit 1 of 2 · BCA Sem 4

Unit 1: Introduction to machine learning

Introduction to Machine Learning notes · PTU syllabus (UGDSE203)

3 min read4 topics8 exam questions
On this page
  1. Unit summary
  2. What is machine learning?
  3. Types of machine learning
  4. Training, validation and testing
  5. Performance metrics
  6. Key terms
  7. Quick revision
  8. Important questions

Unit summary

Machine learning (ML) lets computers learn patterns from data instead of being explicitly programmed. This unit defines ML, traces its history and applications, explains the types of learning, labelled and unlabelled data, regression versus classification, the train-validate-test framework, and how models are evaluated.

After this unit you can

  • Define machine learning and give real-world applications
  • Distinguish supervised, unsupervised, semi-supervised and reinforcement learning
  • Explain training, validation and testing and the problem of overfitting
  • Calculate accuracy, precision, recall and F1 score from a confusion matrix

PTU syllabus topics

  • Definition
  • history and applications of machine learning
  • types of ML (supervised, unsupervised, semi-supervised, reinforcement)
  • labeled/unlabeled datasets
  • regression vs classification
  • training/validation/testing framework
  • performance metrics — confusion matrix
  • accuracy
  • precision
  • recall
  • F1 score
  • AUC
Key termsConfusion matrix metrics
Accuracy
(TP + TN) / total
Precision
TP / (TP + FP): how many predicted positives were right
Recall
TP / (TP + FN): how many actual positives were found
F1 score
Harmonic mean of precision and recall
1

Topic 1

What is machine learning?

Tom Mitchell's definition: a program learns from experience E with respect to task T and performance measure P if its performance on T, measured by P, improves with E.

Example

Spam filter — T: classify emails; E: emails labelled spam or not; P: percentage classified correctly.

History in brief: perceptron (1958), decision trees and backpropagation (1980s), SVMs and random forests (1990s–2000s), deep learning breakthroughs (2012 onward), and today's large language models. Applications: recommendations (Netflix, Amazon), fraud detection, medical diagnosis, speech recognition, self-driving cars and price prediction.

2

Topic 2

Types of machine learning

ComparisonTypes of learning
Data used
Example

Supervised

Labelled data (input + correct output)

Predict house prices, classify emails

Unsupervised

Unlabelled data

Customer segmentation, clustering

Semi-supervised

A little labelled + lots of unlabelled

Photo tagging with few labels

Reinforcement

Rewards and penalties from an environment

Game-playing agents, robotics

Regression predicts a continuous number (price, temperature); classification predicts a category (spam/not spam, pass/fail).

3

Topic 3

Training, validation and testing

ProcessThe ML workflow
  1. 1

    Collect and clean data

  2. 2

    Split data

    e.g. 70% train, 15% validation, 15% test

  3. 3

    Train the model

    On the training set

  4. 4

    Tune hyperparameters

    Using the validation set

  5. 5

    Evaluate once

    On the unseen test set

  6. 6

    Deploy and monitor

  • Overfitting: the model memorises training data and performs poorly on new data (high variance).
  • Underfitting: the model is too simple to capture the pattern (high bias).
  • Cross-validation (k-fold) gives a more reliable estimate by rotating the validation fold.
4

Topic 4

Performance metrics

Predicted positivePredicted negative
Actual positiveTrue Positive (TP)False Negative (FN)
Actual negativeFalse Positive (FP)True Negative (TN)
Key formulasClassification metrics
  • Accuracy

    (TP + TN) / total

  • Precision

    TP / (TP + FP)

    Of predicted positives, how many are right

  • Recall (sensitivity)

    TP / (TP + FN)

    Of actual positives, how many were found

  • F1 score

    2 × Precision × Recall / (Precision + Recall)

  • AUC

    Area under the ROC curve; 1.0 is perfect, 0.5 is random

Example

TP = 40, FP = 10, FN = 20, TN = 30. Accuracy = 70/100 = 70%; precision = 40/50 = 80%; recall = 40/60 ≈ 67%.

Exam tip

Accuracy misleads on imbalanced data — a model that always says "no fraud" is 99% accurate if fraud is 1%. Use precision, recall and F1.

Key terms

Machine learning
Systems that improve at tasks through experience
Labelled data
Data with known correct outputs
Overfitting
Learning noise in training data, failing on new data
Confusion matrix
A table of TP, FP, FN and TN
F1 score
The harmonic mean of precision and recall

Quick revision

  • Supervised (labels), unsupervised (no labels), semi-supervised, reinforcement (rewards).
  • Regression → number; classification → category.
  • Train / validate / test; beware overfitting.
  • Precision = TP/(TP + FP); recall = TP/(TP + FN).

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.Define machine learning with Mitchell's definition.
  2. Q2.Differentiate between supervised and unsupervised learning.
  3. Q3.Differentiate between regression and classification.
  4. Q4.What is overfitting?
  5. Q5.Define precision and recall.

Long-answer questions

  1. Q1.Explain the types of machine learning with examples.
  2. Q2.Explain the training, validation and testing framework and the problems of overfitting and underfitting.
  3. Q3.Explain the confusion matrix and calculate accuracy, precision, recall and F1 for a given example.

Stuck on this unit?

Message SBS on WhatsApp for help with Introduction to Machine Learning, or to ask about studying BCA at Synetic.

WhatsApp us