Unit 1: Introduction to machine learning
Introduction to Machine Learning notes · PTU syllabus (UGDSE203)
On this page
Unit summary
Machine learning (ML) lets computers learn patterns from data instead of being explicitly programmed. This unit defines ML, traces its history and applications, explains the types of learning, labelled and unlabelled data, regression versus classification, the train-validate-test framework, and how models are evaluated.
After this unit you can
- Define machine learning and give real-world applications
- Distinguish supervised, unsupervised, semi-supervised and reinforcement learning
- Explain training, validation and testing and the problem of overfitting
- Calculate accuracy, precision, recall and F1 score from a confusion matrix
PTU syllabus topics
- Definition
- history and applications of machine learning
- types of ML (supervised, unsupervised, semi-supervised, reinforcement)
- labeled/unlabeled datasets
- regression vs classification
- training/validation/testing framework
- performance metrics — confusion matrix
- accuracy
- precision
- recall
- F1 score
- AUC
- Accuracy
- (TP + TN) / total
- Precision
- TP / (TP + FP): how many predicted positives were right
- Recall
- TP / (TP + FN): how many actual positives were found
- F1 score
- Harmonic mean of precision and recall
Topic 1
What is machine learning?
Tom Mitchell's definition: a program learns from experience E with respect to task T and performance measure P if its performance on T, measured by P, improves with E.
Example
Spam filter — T: classify emails; E: emails labelled spam or not; P: percentage classified correctly.
History in brief: perceptron (1958), decision trees and backpropagation (1980s), SVMs and random forests (1990s–2000s), deep learning breakthroughs (2012 onward), and today's large language models. Applications: recommendations (Netflix, Amazon), fraud detection, medical diagnosis, speech recognition, self-driving cars and price prediction.
Topic 2
Types of machine learning
Supervised
Labelled data (input + correct output)
Predict house prices, classify emails
Unsupervised
Unlabelled data
Customer segmentation, clustering
Semi-supervised
A little labelled + lots of unlabelled
Photo tagging with few labels
Reinforcement
Rewards and penalties from an environment
Game-playing agents, robotics
Regression predicts a continuous number (price, temperature); classification predicts a category (spam/not spam, pass/fail).
Topic 3
Training, validation and testing
- 1
Collect and clean data
- 2
Split data
e.g. 70% train, 15% validation, 15% test
- 3
Train the model
On the training set
- 4
Tune hyperparameters
Using the validation set
- 5
Evaluate once
On the unseen test set
- 6
Deploy and monitor
- Overfitting: the model memorises training data and performs poorly on new data (high variance).
- Underfitting: the model is too simple to capture the pattern (high bias).
- Cross-validation (k-fold) gives a more reliable estimate by rotating the validation fold.
Topic 4
Performance metrics
| Predicted positive | Predicted negative | |
|---|---|---|
| Actual positive | True Positive (TP) | False Negative (FN) |
| Actual negative | False Positive (FP) | True Negative (TN) |
Accuracy
(TP + TN) / total
Precision
TP / (TP + FP)
Of predicted positives, how many are right
Recall (sensitivity)
TP / (TP + FN)
Of actual positives, how many were found
F1 score
2 × Precision × Recall / (Precision + Recall)
AUC
Area under the ROC curve; 1.0 is perfect, 0.5 is random
Example
TP = 40, FP = 10, FN = 20, TN = 30. Accuracy = 70/100 = 70%; precision = 40/50 = 80%; recall = 40/60 ≈ 67%.
Exam tip
Accuracy misleads on imbalanced data — a model that always says "no fraud" is 99% accurate if fraud is 1%. Use precision, recall and F1.
Key terms
- Machine learning
- Systems that improve at tasks through experience
- Labelled data
- Data with known correct outputs
- Overfitting
- Learning noise in training data, failing on new data
- Confusion matrix
- A table of TP, FP, FN and TN
- F1 score
- The harmonic mean of precision and recall
Quick revision
- Supervised (labels), unsupervised (no labels), semi-supervised, reinforcement (rewards).
- Regression → number; classification → category.
- Train / validate / test; beware overfitting.
- Precision = TP/(TP + FP); recall = TP/(TP + FN).
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.Define machine learning with Mitchell's definition.
- Q2.Differentiate between supervised and unsupervised learning.
- Q3.Differentiate between regression and classification.
- Q4.What is overfitting?
- Q5.Define precision and recall.
Long-answer questions
- Q1.Explain the types of machine learning with examples.
- Q2.Explain the training, validation and testing framework and the problems of overfitting and underfitting.
- Q3.Explain the confusion matrix and calculate accuracy, precision, recall and F1 for a given example.
Stuck on this unit?
Message SBS on WhatsApp for help with Introduction to Machine Learning, or to ask about studying BCA at Synetic.
