Unit 1 of 1 · M.Sc IT Sem 4

Unit 1: ML model implementation

Machine Learning Laboratory notes · PTU syllabus (PGCA1946)

3 min read7 topics10 exam questions
On this page
  1. Unit summary
  2. The ML workflow
  3. Linear regression
  4. Logistic regression
  5. K-nearest neighbours
  6. K-means clustering
  7. Support vector machines
  8. Recording and comparing results
  9. Key terms
  10. Quick revision
  11. Important questions

Unit summary

This lab designs and evaluates machine learning models — linear regression, logistic regression, k-nearest neighbours, k-means clustering and support vector machines — on standard datasets in Python (R and MATLAB equivalents follow the same steps).

After this unit you can

  • Prepare data and split it for training and testing
  • Train and evaluate regression and classification models
  • Cluster data with k-means and choose k
  • Compare models with appropriate metrics

PTU syllabus topics

  • Data model design and evaluation using Linear Regression
  • Logistic Regression
  • K-Nearest Neighbors
  • K-Means Clustering and Support Vector Machines
  • using standard or assumed datasets in Python/R/MATLAB
ProcessML lab workflow in Python
  1. 1

    Load data

    pandas.read_csv

  2. 2

    Explore and clean

  3. 3

    Split train and test

    train_test_split

  4. 4

    Fit the model

    LinearRegression, KNN, SVM

  5. 5

    Evaluate

    Accuracy, confusion matrix, RMSE

  6. 6

    Visualise results

1

Topic 1

The ML workflow

ProcessLab workflow
  1. 1

    Load the dataset (scikit-learn, CSV or UCI)

  2. 2

    Explore — shape, describe(), missing values, plots

  3. 3

    Preprocess — impute, encode, scale

  4. 4

    Split train and test (e.g., 80/20, stratified)

  5. 5

    Train the model

  6. 6

    Predict and evaluate

  7. 7

    Tune with cross-validation

  8. 8

    Record results in a table

pythonimport pandas as pd, numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
2

Topic 2

Linear regression

pythonfrom sklearn.datasets import fetch_california_housing
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, r2_score
X, y = fetch_california_housing(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
m = LinearRegression().fit(Xtr, ytr); p = m.predict(Xte)
print("RMSE", mean_squared_error(yte, p) ** 0.5, "R2", r2_score(yte, p))
print(dict(zip(fetch_california_housing().feature_names, m.coef_.round(3))))
  • Plot predicted vs actual and residuals; a funnel shape suggests non-constant variance.
3

Topic 3

Logistic regression

pythonfrom sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report, confusion_matrix, roc_auc_score
X, y = load_breast_cancer(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
sc = StandardScaler().fit(Xtr); Xtr, Xte = sc.transform(Xtr), sc.transform(Xte)
lr = LogisticRegression(max_iter=1000).fit(Xtr, ytr)
print(confusion_matrix(yte, lr.predict(Xte))); print(classification_report(yte, lr.predict(Xte)))
print("AUC", roc_auc_score(yte, lr.predict_proba(Xte)[:, 1]))
4

Topic 4

K-nearest neighbours

pythonfrom sklearn.datasets import load_iris
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_iris(return_X_y=True)
for k in (1, 3, 5, 7, 9):
    print(k, cross_val_score(KNeighborsClassifier(k), X, y, cv=5).mean().round(3))
  • Pick k with the best cross-validated accuracy; always scale features for distance-based models.
5

Topic 5

K-means clustering

pythonfrom sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
import matplotlib.pyplot as plt
X, _ = load_iris(return_X_y=True); X = StandardScaler().fit_transform(X)
wss = []
for k in range(1, 9):
    km = KMeans(n_clusters=k, n_init=10, random_state=0).fit(X); wss.append(km.inertia_)
    if k > 1: print(k, round(silhouette_score(X, km.labels_), 3))
plt.plot(range(1, 9), wss, marker="o"); plt.xlabel("k"); plt.ylabel("WSS"); plt.show()   # elbow
6

Topic 6

Support vector machines

pythonfrom sklearn.svm import SVC
from sklearn.model_selection import GridSearchCV
X, y = load_breast_cancer(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
sc = StandardScaler().fit(Xtr)
g = GridSearchCV(SVC(), {"kernel": ["linear", "rbf"], "C": [0.1, 1, 10], "gamma": ["scale", 0.01]}, cv=5)
g.fit(sc.transform(Xtr), ytr)
print(g.best_params_, g.score(sc.transform(Xte), yte))
7

Topic 7

Recording and comparing results

ComparisonModel comparison sheet
Metric to report
Notes

Linear regression

RMSE, MAE, R²

Check residual plots

Logistic regression

Accuracy, precision, recall, F1, AUC

Use a confusion matrix

k-NN

Cross-validated accuracy vs k

Scale features

k-means

WSS (elbow), silhouette

No labels needed

SVM

Accuracy, F1 with best C, gamma, kernel

Grid search with CV

  • R: lm(), glm(family = binomial), class::knn(), kmeans(), e1071::svm(). MATLAB: fitlm, fitglm, fitcknn, kmeans, fitcsvm.

Key terms

Train–test split
Holding out data to estimate generalisation
Stratify
Keeping class proportions equal across splits
Cross-validation
Averaging performance over k folds
Grid search
Trying combinations of hyperparameters
Inertia
Within-cluster sum of squares in k-means

Quick revision

  • Load, explore, preprocess, split, train, evaluate, tune.
  • Linear regression: RMSE, R²; logistic: confusion matrix, F1, AUC.
  • k-NN: choose k by CV; k-means: elbow and silhouette.
  • SVM: kernels, C, gamma, grid search; compare in a table.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.Why scale features before k-NN and SVM?
  2. Q2.What does R² measure?
  3. Q3.What is a confusion matrix?
  4. Q4.How do you choose k in k-means?
  5. Q5.What does GridSearchCV do?
  6. Q6.Why stratify the split?

Long-answer questions

  1. Q1.Build and evaluate a linear regression model on a dataset.
  2. Q2.Build a logistic regression classifier and interpret its metrics.
  3. Q3.Apply k-means clustering and choose the number of clusters.
  4. Q4.Train an SVM with hyperparameter tuning and compare it with k-NN.

Stuck on this unit?

Message SBS on WhatsApp for help with Machine Learning Laboratory, or to ask about studying M.Sc IT at Synetic.

WhatsApp us