Unit 4 of 4 · B.Sc IT Sem 5

Unit 4: Prediction, clustering and visualization

Data Warehousing & Mining notes · PTU syllabus (BSIT504/BSBC501)

3 min read7 topics10 exam questions
On this page
  1. Unit summary
  2. Prediction
  3. Classifier accuracy
  4. Holdout, cross-validation and bootstrap
  5. Bagging and boosting
  6. Clustering and the classification of clustering algorithms
  7. Selecting the right data mining technique
  8. Data visualisation
  9. Key terms
  10. Quick revision
  11. Important questions

Unit summary

Beyond classifying, data mining predicts values, measures accuracy, groups similar records and presents findings visually. This unit covers prediction, classifier accuracy, cross-validation, bootstrap, boosting and bagging, the classification of clustering algorithms, selecting the right data mining technique and data visualisation.

After this unit you can

  • Explain prediction techniques
  • Estimate classifier accuracy with holdout, cross-validation and bootstrap
  • Explain bagging and boosting
  • Classify clustering methods, choose techniques and visualise data

PTU syllabus topics

  • Prediction techniques
  • classifier accuracy
  • cross-validation
  • bootstrap
  • boosting
  • bagging
  • clustering algorithm classification
  • selecting the right data mining technique
  • data visualization
ClassificationData mining techniques
Data mining
  • Classification

    Decision trees, Naive Bayes

  • Prediction

    Regression

  • Clustering

    k-means, hierarchical

  • Association

    Apriori, market basket

  • Ensembles

    Bagging, boosting

1

Topic 1

Prediction

  • Prediction estimates a continuous value (sales, price) rather than a class label.
Key termsPrediction techniques
Linear regression
y = a + bx fitted by least squares
Multiple regression
Several predictors
Non-linear regression
Polynomial or transformed variables
Regression trees
Leaves hold numeric values
Neural networks and k-NN
Learn complex patterns

Example

Advertising spend (₹ lakh) and sales: fitted line sales = 20 + 3.5 × spend; at spend 10, predicted sales = 55.

2

Topic 2

Classifier accuracy

Predicted yesPredicted no
Actual yesTrue positives (TP)False negatives (FN)
Actual noFalse positives (FP)True negatives (TN)
Key formulasAccuracy measures
  • Accuracy

    (TP + TN) ÷ total

  • Error rate

    1 − accuracy

  • Precision

    TP ÷ (TP + FP)

  • Recall (sensitivity)

    TP ÷ (TP + FN)

  • Specificity

    TN ÷ (TN + FP)

  • F1 score

    2 × precision × recall ÷ (precision + recall)

Example

TP = 90, FN = 10, FP = 30, TN = 870: accuracy = 960 ÷ 1,000 = 96%; precision = 90 ÷ 120 = 75%; recall = 90 ÷ 100 = 90%.

3

Topic 3

Holdout, cross-validation and bootstrap

ComparisonAccuracy estimation methods
How it works
Note

Holdout

Split data, e.g., two-thirds training and one-third testing

Simple; result depends on the split

Random subsampling

Repeat holdout k times and average

More stable

k-fold cross-validation

Divide into k folds; each fold tests once while k − 1 train; average

10-fold is standard; leave-one-out when k = n

Bootstrap

Sample n records with replacement for training; unused records test

About 63.2% of records appear in training (.632 bootstrap)

4

Topic 4

Bagging and boosting

ComparisonEnsemble methods
Bagging
Boosting

Full form

Bootstrap aggregating

—

Training sets

Bootstrap samples drawn independently

Sequential; weights increased for misclassified records

Combining

Equal-weight majority vote

Weighted vote by each model's accuracy

Effect

Reduces variance; robust to noise

Reduces bias; may overfit noisy data

Example

Random forest

AdaBoost, gradient boosting (XGBoost)

5

Topic 5

Clustering and the classification of clustering algorithms

  • Clustering: grouping objects so that those in a cluster are similar and different from other clusters — unsupervised learning; uses distance measures such as Euclidean and Manhattan distance.
ClassificationClustering methods
Clustering
  • Partitioning

    k-means, k-medoids (PAM)

  • Hierarchical

    Agglomerative (bottom-up), divisive (top-down); dendrogram

  • Density-based

    DBSCAN, OPTICS — arbitrary shapes, handles noise

  • Grid-based

    STING, CLIQUE — fast, on a grid of cells

  • Model-based

    EM, self-organising maps

Processk-means algorithm
  1. 1Choose k
  2. 2Pick k initial centroids
  3. 3Assign each point to the nearest centroid
  4. 4Recompute centroids as cluster means
  5. 5Repeat until assignments stop changing

Example

Points 2, 4, 10, 12, 3, 20, 30, 11, 25 with k = 2 and initial means 2 and 4 converge to clusters {2, 3, 4, 10, 11, 12} (mean 7) and {20, 25, 30} (mean 25).

  • Applications: customer segmentation, document grouping, image segmentation, anomaly detection.
6

Topic 6

Selecting the right data mining technique

ComparisonChoosing a technique
Business question
Technique

Which items go together?

Basket composition

Association rules

Will this customer default?

Known classes

Classification — decision tree, naive Bayes

How much will we sell?

Numeric target

Prediction — regression

What natural groups exist?

No labels

Clustering

Which transactions are unusual?

Rare events

Outlier analysis

  • Other factors: data size and type, need for explainable results (decision trees) versus accuracy (ensembles), speed, and the skills and tools available.
7

Topic 7

Data visualisation

  • Visualisation presents data and mining results graphically to reveal patterns, support exploration and communicate insights.
Key termsVisualisation techniques
Pixel-oriented
Each value a coloured pixel
Geometric projection
Scatter plot matrices, parallel coordinates
Icon-based
Chernoff faces, stick figures
Hierarchical
Tree maps, cone trees
Charts and dashboards
Bar, line, pie, heat maps in Power BI or Tableau
Mining result views
Decision tree diagrams, rule graphs, cluster plots

Key terms

Prediction
Estimating a continuous value
Cross-validation
Rotating training and test folds to estimate accuracy
Bagging
Ensemble of models trained on bootstrap samples
Boosting
Sequential ensemble focusing on misclassified records
k-means
Partitioning method assigning points to the nearest centroid

Quick revision

  • Regression and other prediction methods.
  • Confusion matrix; accuracy, precision, recall, F1.
  • Holdout, random subsampling, k-fold, bootstrap.
  • Bagging vs boosting.
  • Partitioning, hierarchical, density, grid, model clustering; k-means; choosing techniques; visualisation types.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.Distinguish classification and prediction.
  2. Q2.Define precision and recall.
  3. Q3.What is 10-fold cross-validation?
  4. Q4.Distinguish bagging and boosting.
  5. Q5.Name the categories of clustering methods.
  6. Q6.Which technique answers "what natural groups exist?"

Long-answer questions

  1. Q1.Explain prediction techniques and measures of classifier accuracy.
  2. Q2.Explain cross-validation, bootstrap, bagging and boosting.
  3. Q3.Explain the classification of clustering methods with the k-means algorithm.
  4. Q4.Explain how to select the right data mining technique and the role of visualisation.

Stuck on this unit?

Message SBS on WhatsApp for help with Data Warehousing & Mining, or to ask about studying B.Sc IT at Synetic.

WhatsApp us