Unit 1: ML model implementation
Machine Learning Laboratory notes · PTU syllabus (PGCA1946)
On this page
Unit summary
This lab designs and evaluates machine learning models — linear regression, logistic regression, k-nearest neighbours, k-means clustering and support vector machines — on standard datasets in Python (R and MATLAB equivalents follow the same steps).
After this unit you can
- Prepare data and split it for training and testing
- Train and evaluate regression and classification models
- Cluster data with k-means and choose k
- Compare models with appropriate metrics
PTU syllabus topics
- Data model design and evaluation using Linear Regression
- Logistic Regression
- K-Nearest Neighbors
- K-Means Clustering and Support Vector Machines
- using standard or assumed datasets in Python/R/MATLAB
- 1
Load data
pandas.read_csv
- 2
Explore and clean
- 3
Split train and test
train_test_split
- 4
Fit the model
LinearRegression, KNN, SVM
- 5
Evaluate
Accuracy, confusion matrix, RMSE
- 6
Visualise results
Topic 1
The ML workflow
- 1
Load the dataset (scikit-learn, CSV or UCI)
- 2
Explore — shape, describe(), missing values, plots
- 3
Preprocess — impute, encode, scale
- 4
Split train and test (e.g., 80/20, stratified)
- 5
Train the model
- 6
Predict and evaluate
- 7
Tune with cross-validation
- 8
Record results in a table
pythonimport pandas as pd, numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScalerTopic 2
Linear regression
pythonfrom sklearn.datasets import fetch_california_housing
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, r2_score
X, y = fetch_california_housing(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
m = LinearRegression().fit(Xtr, ytr); p = m.predict(Xte)
print("RMSE", mean_squared_error(yte, p) ** 0.5, "R2", r2_score(yte, p))
print(dict(zip(fetch_california_housing().feature_names, m.coef_.round(3))))- Plot predicted vs actual and residuals; a funnel shape suggests non-constant variance.
Topic 3
Logistic regression
pythonfrom sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report, confusion_matrix, roc_auc_score
X, y = load_breast_cancer(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
sc = StandardScaler().fit(Xtr); Xtr, Xte = sc.transform(Xtr), sc.transform(Xte)
lr = LogisticRegression(max_iter=1000).fit(Xtr, ytr)
print(confusion_matrix(yte, lr.predict(Xte))); print(classification_report(yte, lr.predict(Xte)))
print("AUC", roc_auc_score(yte, lr.predict_proba(Xte)[:, 1]))Topic 4
K-nearest neighbours
pythonfrom sklearn.datasets import load_iris
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_iris(return_X_y=True)
for k in (1, 3, 5, 7, 9):
print(k, cross_val_score(KNeighborsClassifier(k), X, y, cv=5).mean().round(3))- Pick k with the best cross-validated accuracy; always scale features for distance-based models.
Topic 5
K-means clustering
pythonfrom sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
import matplotlib.pyplot as plt
X, _ = load_iris(return_X_y=True); X = StandardScaler().fit_transform(X)
wss = []
for k in range(1, 9):
km = KMeans(n_clusters=k, n_init=10, random_state=0).fit(X); wss.append(km.inertia_)
if k > 1: print(k, round(silhouette_score(X, km.labels_), 3))
plt.plot(range(1, 9), wss, marker="o"); plt.xlabel("k"); plt.ylabel("WSS"); plt.show() # elbowTopic 6
Support vector machines
pythonfrom sklearn.svm import SVC
from sklearn.model_selection import GridSearchCV
X, y = load_breast_cancer(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
sc = StandardScaler().fit(Xtr)
g = GridSearchCV(SVC(), {"kernel": ["linear", "rbf"], "C": [0.1, 1, 10], "gamma": ["scale", 0.01]}, cv=5)
g.fit(sc.transform(Xtr), ytr)
print(g.best_params_, g.score(sc.transform(Xte), yte))Topic 7
Recording and comparing results
Linear regression
RMSE, MAE, R²
Check residual plots
Logistic regression
Accuracy, precision, recall, F1, AUC
Use a confusion matrix
k-NN
Cross-validated accuracy vs k
Scale features
k-means
WSS (elbow), silhouette
No labels needed
SVM
Accuracy, F1 with best C, gamma, kernel
Grid search with CV
- R: lm(), glm(family = binomial), class::knn(), kmeans(), e1071::svm(). MATLAB: fitlm, fitglm, fitcknn, kmeans, fitcsvm.
Key terms
- Train–test split
- Holding out data to estimate generalisation
- Stratify
- Keeping class proportions equal across splits
- Cross-validation
- Averaging performance over k folds
- Grid search
- Trying combinations of hyperparameters
- Inertia
- Within-cluster sum of squares in k-means
Quick revision
- Load, explore, preprocess, split, train, evaluate, tune.
- Linear regression: RMSE, R²; logistic: confusion matrix, F1, AUC.
- k-NN: choose k by CV; k-means: elbow and silhouette.
- SVM: kernels, C, gamma, grid search; compare in a table.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.Why scale features before k-NN and SVM?
- Q2.What does R² measure?
- Q3.What is a confusion matrix?
- Q4.How do you choose k in k-means?
- Q5.What does GridSearchCV do?
- Q6.Why stratify the split?
Long-answer questions
- Q1.Build and evaluate a linear regression model on a dataset.
- Q2.Build a logistic regression classifier and interpret its metrics.
- Q3.Apply k-means clustering and choose the number of clusters.
- Q4.Train an SVM with hyperparameter tuning and compare it with k-NN.
Stuck on this unit?
Message SBS on WhatsApp for help with Machine Learning Laboratory, or to ask about studying M.Sc IT at Synetic.
