Unit 4: Prediction, clustering and visualization
Data Warehousing & Mining notes · PTU syllabus (BSIT504/BSBC501)
On this page
Unit summary
Beyond classifying, data mining predicts values, measures accuracy, groups similar records and presents findings visually. This unit covers prediction, classifier accuracy, cross-validation, bootstrap, boosting and bagging, the classification of clustering algorithms, selecting the right data mining technique and data visualisation.
After this unit you can
- Explain prediction techniques
- Estimate classifier accuracy with holdout, cross-validation and bootstrap
- Explain bagging and boosting
- Classify clustering methods, choose techniques and visualise data
PTU syllabus topics
- Prediction techniques
- classifier accuracy
- cross-validation
- bootstrap
- boosting
- bagging
- clustering algorithm classification
- selecting the right data mining technique
- data visualization
Classification
Decision trees, Naive Bayes
Prediction
Regression
Clustering
k-means, hierarchical
Association
Apriori, market basket
Ensembles
Bagging, boosting
Topic 1
Prediction
- Prediction estimates a continuous value (sales, price) rather than a class label.
- Linear regression
- y = a + bx fitted by least squares
- Multiple regression
- Several predictors
- Non-linear regression
- Polynomial or transformed variables
- Regression trees
- Leaves hold numeric values
- Neural networks and k-NN
- Learn complex patterns
Example
Advertising spend (₹ lakh) and sales: fitted line sales = 20 + 3.5 × spend; at spend 10, predicted sales = 55.
Topic 2
Classifier accuracy
| Predicted yes | Predicted no | |
|---|---|---|
| Actual yes | True positives (TP) | False negatives (FN) |
| Actual no | False positives (FP) | True negatives (TN) |
Accuracy
(TP + TN) ÷ total
Error rate
1 − accuracy
Precision
TP ÷ (TP + FP)
Recall (sensitivity)
TP ÷ (TP + FN)
Specificity
TN ÷ (TN + FP)
F1 score
2 × precision × recall ÷ (precision + recall)
Example
TP = 90, FN = 10, FP = 30, TN = 870: accuracy = 960 ÷ 1,000 = 96%; precision = 90 ÷ 120 = 75%; recall = 90 ÷ 100 = 90%.
Topic 3
Holdout, cross-validation and bootstrap
Holdout
Split data, e.g., two-thirds training and one-third testing
Simple; result depends on the split
Random subsampling
Repeat holdout k times and average
More stable
k-fold cross-validation
Divide into k folds; each fold tests once while k − 1 train; average
10-fold is standard; leave-one-out when k = n
Bootstrap
Sample n records with replacement for training; unused records test
About 63.2% of records appear in training (.632 bootstrap)
Topic 4
Bagging and boosting
Full form
Bootstrap aggregating
—
Training sets
Bootstrap samples drawn independently
Sequential; weights increased for misclassified records
Combining
Equal-weight majority vote
Weighted vote by each model's accuracy
Effect
Reduces variance; robust to noise
Reduces bias; may overfit noisy data
Example
Random forest
AdaBoost, gradient boosting (XGBoost)
Topic 5
Clustering and the classification of clustering algorithms
- Clustering: grouping objects so that those in a cluster are similar and different from other clusters — unsupervised learning; uses distance measures such as Euclidean and Manhattan distance.
Partitioning
k-means, k-medoids (PAM)
Hierarchical
Agglomerative (bottom-up), divisive (top-down); dendrogram
Density-based
DBSCAN, OPTICS — arbitrary shapes, handles noise
Grid-based
STING, CLIQUE — fast, on a grid of cells
Model-based
EM, self-organising maps
- 1Choose k
- 2Pick k initial centroids
- 3Assign each point to the nearest centroid
- 4Recompute centroids as cluster means
- 5Repeat until assignments stop changing
Example
Points 2, 4, 10, 12, 3, 20, 30, 11, 25 with k = 2 and initial means 2 and 4 converge to clusters {2, 3, 4, 10, 11, 12} (mean 7) and {20, 25, 30} (mean 25).
- Applications: customer segmentation, document grouping, image segmentation, anomaly detection.
Topic 6
Selecting the right data mining technique
Which items go together?
Basket composition
Association rules
Will this customer default?
Known classes
Classification — decision tree, naive Bayes
How much will we sell?
Numeric target
Prediction — regression
What natural groups exist?
No labels
Clustering
Which transactions are unusual?
Rare events
Outlier analysis
- Other factors: data size and type, need for explainable results (decision trees) versus accuracy (ensembles), speed, and the skills and tools available.
Topic 7
Data visualisation
- Visualisation presents data and mining results graphically to reveal patterns, support exploration and communicate insights.
- Pixel-oriented
- Each value a coloured pixel
- Geometric projection
- Scatter plot matrices, parallel coordinates
- Icon-based
- Chernoff faces, stick figures
- Hierarchical
- Tree maps, cone trees
- Charts and dashboards
- Bar, line, pie, heat maps in Power BI or Tableau
- Mining result views
- Decision tree diagrams, rule graphs, cluster plots
Key terms
- Prediction
- Estimating a continuous value
- Cross-validation
- Rotating training and test folds to estimate accuracy
- Bagging
- Ensemble of models trained on bootstrap samples
- Boosting
- Sequential ensemble focusing on misclassified records
- k-means
- Partitioning method assigning points to the nearest centroid
Quick revision
- Regression and other prediction methods.
- Confusion matrix; accuracy, precision, recall, F1.
- Holdout, random subsampling, k-fold, bootstrap.
- Bagging vs boosting.
- Partitioning, hierarchical, density, grid, model clustering; k-means; choosing techniques; visualisation types.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.Distinguish classification and prediction.
- Q2.Define precision and recall.
- Q3.What is 10-fold cross-validation?
- Q4.Distinguish bagging and boosting.
- Q5.Name the categories of clustering methods.
- Q6.Which technique answers "what natural groups exist?"
Long-answer questions
- Q1.Explain prediction techniques and measures of classifier accuracy.
- Q2.Explain cross-validation, bootstrap, bagging and boosting.
- Q3.Explain the classification of clustering methods with the k-means algorithm.
- Q4.Explain how to select the right data mining technique and the role of visualisation.
Stuck on this unit?
Message SBS on WhatsApp for help with Data Warehousing & Mining, or to ask about studying B.Sc IT at Synetic.
