Unit 4: Evaluation and validation
Data Sciences Using R notes · PTU syllabus (MBA 962-18)
On this page
Unit summary
A model is only useful if its performance on new data is known and significant. This unit covers methods for estimating classifier performance — cross-validation, the holdout method, the bootstrap method and the confusion matrix — assessing the statistical significance of data mining results, and advanced topics: scalable machine learning, big data techniques, stream data mining and social network analysis.
After this unit you can
- Estimate classifier performance with holdout, cross-validation and bootstrap
- Interpret the confusion matrix and related metrics
- Assess the statistical significance of data mining results
- Explain scalable ML, big data techniques, stream mining and social networks
PTU syllabus topics
- Methods for estimating classifier performance — cross-validation
- holdout method
- bootstrap method
- confusion matrix
- assessing statistical significance of data mining results
- advanced topics (scalable ML, big data techniques, stream data mining, social networks)
Holdout
Single train/test split, e.g. 70:30
Fast; result depends on the split
k-fold cross-validation
Train on k − 1 folds, test on 1, repeat
Reliable; standard choice
Bootstrap
Resample with replacement
Good for small data
Confusion matrix
TP, FP, TN, FN counts
Basis for accuracy, precision, recall
Topic 1
Holdout method
- Split data into training (e.g., 70%) and test (30%) sets; build the model on training data and evaluate on test data.
- Stratified sampling keeps class proportions similar in both sets.
- Limitation: results depend on one random split; wasteful with small data.
- In R: caret::createDataPartition(y, p = 0.7).
Topic 2
Cross-validation
- 1Split data into k equal folds
- 2Train on k − 1 folds
- 3Test on the remaining fold
- 4Repeat k times so each fold is tested once
- 5Average the performance
- Common choices: 5-fold or 10-fold; leave-one-out (k = n) for very small data; repeated cross-validation for stability.
- In R: caret::trainControl(method = "cv", number = 10).
Topic 3
Bootstrap method
- Draw samples with replacement of the same size as the data; train on the bootstrap sample and test on the left-out (out-of-bag) observations — about 36.8% of cases are left out each time.
- .632 bootstrap: combines training and out-of-bag error to reduce bias.
Topic 4
Confusion matrix
| Predicted positive | Predicted negative | |
|---|---|---|
| Actual positive | True positive (TP) | False negative (FN) |
| Actual negative | False positive (FP) | True negative (TN) |
Accuracy
(TP + TN) ÷ total
Precision
TP ÷ (TP + FP)
Recall (sensitivity)
TP ÷ (TP + FN)
Specificity
TN ÷ (TN + FP)
F1 score
2 × precision × recall ÷ (precision + recall)
Example
Of 1,000 customers, 100 churn. The model flags 120: 80 actual churners (TP), 40 non-churners (FP). Precision = 80 ÷ 120 = 67%; recall = 80 ÷ 100 = 80%; accuracy = (80 + 860) ÷ 1,000 = 94%.
- ROC curve and AUC: trade-off between sensitivity and false positive rate across cut-offs (pROC package).
- Imbalanced data: accuracy misleads — use precision, recall, F1, AUC; rebalance with oversampling (SMOTE) or class weights.
Topic 5
Statistical significance of data mining results
- Why: with many models and variables, some patterns appear by chance (data snooping).
- Comparing two classifiers: paired t-test on cross-validation folds, McNemar's test on the same test set.
- Confidence intervals for accuracy; permutation tests (shuffle labels to see whether performance beats chance).
- Multiple comparisons: adjust significance levels (Bonferroni) when testing many patterns or rules.
Topic 6
Advanced topics
- Scalable machine learning: distributed training (Spark MLlib, sparklyr in R), stochastic gradient descent, online learning, GPUs.
- Big data techniques: MapReduce, Hadoop and Spark, NoSQL stores, data lakes, sampling and dimensionality reduction for very large data.
- Stream data mining: analysing continuous data in real time (sensor feeds, clickstreams, transactions) with one-pass algorithms, sliding windows and concept-drift detection — e.g., real-time fraud alerts.
- Social network analysis: nodes and edges; centrality measures (degree, betweenness, closeness, eigenvector), community detection, influence and diffusion — igraph package in R.
Key terms
- Holdout method
- Single train–test split for evaluation
- k-fold cross-validation
- Rotating training and test folds
- Recall
- Share of actual positives correctly identified
- Permutation test
- Significance test by shuffling labels
- Concept drift
- Change over time in the patterns a model learned
Quick revision
- Holdout (stratified split); k-fold and leave-one-out CV; bootstrap and .632.
- Confusion matrix: TP, FP, FN, TN; accuracy, precision, recall, specificity, F1; ROC–AUC.
- Imbalanced data: SMOTE, class weights.
- Significance: paired t-test, McNemar, permutation, Bonferroni.
- Scalable ML, Spark, stream mining, social network centrality.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.What is the holdout method?
- Q2.Explain 10-fold cross-validation.
- Q3.What fraction of data is out-of-bag in a bootstrap sample?
- Q4.Define precision and recall.
- Q5.Why is accuracy misleading for imbalanced data?
- Q6.What is stream data mining?
Long-answer questions
- Q1.Explain holdout, cross-validation and bootstrap methods of estimating classifier performance.
- Q2.Explain the confusion matrix and the metrics derived from it with an example.
- Q3.Discuss how the statistical significance of data mining results is assessed.
- Q4.Discuss scalable machine learning, big data techniques, stream mining and social network analysis.
Stuck on this unit?
Message SBS on WhatsApp for help with Data Sciences Using R, or to ask about studying MBA at Synetic.
