Unit 1: ML model implementation in Python
Introduction to Machine Learning Laboratory notes · PTU syllabus (UGDSE204)
On this page
Unit summary
This lab implements the main ML algorithms in Python with scikit-learn and visualises their results: regression lines, decision boundaries, trees, Naive Bayes, random forests, SVMs, three clustering methods, PCA, perceptrons and AdaBoost.
After this unit you can
- Train and evaluate regression and classification models
- Visualise regression lines, decision boundaries and dendrograms
- Apply K-means, hierarchical clustering, DBSCAN and PCA
- Demonstrate perceptrons and boosting
PTU syllabus topics
- Linear regression with regression line visualization
- logistic regression with decision boundary
- decision tree (ID3/CART) classifier
- Naive Bayes classifier
- random forest classifier
- SVM for linearly separable classes
- K-Means clustering with visualization
- hierarchical clustering with dendrogram
- DBSCAN clustering
- PCA with classifier performance comparison
- single-layer perceptron for AND/OR/XOR
- AdaBoost boosting demonstration
- 1Load data
- 2Split
Training and test sets
- 3Train
Fit the model
- 4Evaluate
Accuracy, precision, recall
- 5Visualise
Decision boundary, clusters or tree
Topic 1
The standard scikit-learn pattern
pythonfrom sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, confusion_matrix
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = SomeModel() # e.g. LogisticRegression()
model.fit(X_train, y_train)
pred = model.predict(X_test)
print(accuracy_score(y_test, pred), confusion_matrix(y_test, pred))| Experiment | scikit-learn class |
|---|---|
| Linear regression | LinearRegression |
| Logistic regression | LogisticRegression |
| Decision tree (ID3/CART) | DecisionTreeClassifier(criterion="entropy" or "gini") |
| Naive Bayes | GaussianNB |
| Random forest | RandomForestClassifier |
| SVM (linear) | SVC(kernel="linear") |
| K-means | KMeans(n_clusters=k) |
| Hierarchical | AgglomerativeClustering; scipy dendrogram |
| DBSCAN | DBSCAN(eps, min_samples) |
| PCA | PCA(n_components=2) |
| Boosting | AdaBoostClassifier |
Topic 2
Visualising models
pythonimport matplotlib.pyplot as plt
plt.scatter(X, y); plt.plot(X, model.predict(X), color="red") # regression line
from scipy.cluster.hierarchy import dendrogram, linkage
dendrogram(linkage(X, method="ward")); plt.show()For a decision boundary, predict on a grid of points (np.meshgrid) and draw it with plt.contourf.
Topic 3
Perceptron for AND, OR and XOR
pythonimport numpy as np
def train(X, t, lr=0.1, epochs=20):
w, b = np.zeros(2), 0
for _ in range(epochs):
for xi, ti in zip(X, t):
y = 1 if xi @ w + b > 0 else 0
w += lr * (ti - y) * xi; b += lr * (ti - y)
return w, b
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]])
train(X, [0, 0, 0, 1]) # AND: learns correctly
train(X, [0, 1, 1, 0]) # XOR: never converges — not linearly separableTopic 4
PCA and boosting
Apply PCA to reduce features to 2, train the same classifier before and after, and compare accuracy and training time. AdaBoost combines many weak learners (shallow trees), giving more weight to misclassified samples each round.
Exam tip
Always set random_state for reproducible results and report accuracy with a confusion matrix — examiners look for both.
Key terms
- train_test_split
- Splits data into training and test sets
- fit / predict
- Train a model / make predictions
- Dendrogram
- A tree diagram of hierarchical clustering
- AdaBoost
- A boosting method combining weak learners
- Decision boundary
- The line or surface separating predicted classes
Quick revision
- fit on training data, evaluate on test data.
- criterion="entropy" ≈ ID3; "gini" ≈ CART.
- Perceptron learns AND/OR, fails on XOR.
- PCA then classify to compare performance.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.What does train_test_split do?
- Q2.Which class implements Naive Bayes in scikit-learn?
- Q3.What do eps and min_samples mean in DBSCAN?
- Q4.Why does the perceptron fail on XOR?
- Q5.What is AdaBoost?
Long-answer questions
- Q1.Implement linear regression and plot the regression line.
- Q2.Implement a decision tree and a random forest on the same data set and compare accuracy.
- Q3.Implement K-means and hierarchical clustering and visualise the clusters and dendrogram.
- Q4.Implement a single-layer perceptron for AND, OR and XOR and explain the results.
Stuck on this unit?
Message SBS on WhatsApp for help with Introduction to Machine Learning Laboratory, or to ask about studying BCA at Synetic.
