Unit 2: Advanced analytical methods
Big Data Analytics notes · PTU syllabus (PGCA1947)
On this page
Unit summary
Analytical methods turn big data into insight. This unit covers k-means clustering use cases and diagnostics, decision trees and their evaluation, Naive Bayes, association rules and Apriori with rule evaluation, finding similarity, and recommendation systems.
After this unit you can
- Apply k-means and diagnose cluster quality
- Build and evaluate decision trees and Naive Bayes
- Evaluate association rules with support, confidence and lift
- Compare recommendation system approaches
PTU syllabus topics
- K-means clustering use cases and diagnostics
- decision tree algorithms and evaluation
- Naive Bayes classification
- association rules and the Apriori algorithm
- evaluation of candidate rules
- finding association and similarity
- collaborative/content-based/knowledge-based/hybrid recommendation systems
Collaborative filtering
Similar users' behaviour
People who bought this also bought
Content-based
Item features
Similar movies by genre
Knowledge-based
Explicit rules and needs
Configure a laptop by requirements
Hybrid
A mix
Netflix-style recommendations
Topic 1
K-means clustering: use cases and diagnostics
- 1Choose k
- 2Initialise centroids
- 3Assign to nearest centroid
- 4Recompute means
- 5Repeat until stable
- WSS (within sum of squares)
- Plot against k; elbow suggests k
- Silhouette
- −1 to 1; higher means well-separated
- Scaling
- Standardise features
- Use cases
- Customer segmentation, image compression, document grouping
Topic 2
Decision trees and their evaluation
Entropy = −Σ p log2 p
Impurity
Gain = Entropy(parent) − weighted Entropy(children)
Split choice
Gini = 1 − Σ p²
CART impurity
- ID3, C4.5, CART
- Common algorithms
- Confusion matrix
- TP, FP, TN, FN
- Accuracy, precision, recall
- Performance
- ROC curve and AUC
- Trade-off across thresholds
- Pruning
- Avoid overfitting
Topic 3
Naive Bayes
P(C given X) ∝ P(C) × Π P(xi given C)
Independence assumption
Laplace smoothing
Add 1 to counts
- Fast, scales well to large text data (spam filtering, sentiment).
Topic 4
Association rules and Apriori
Support(A → B) = P(A and B)
Frequency
Confidence = Support(A and B) / Support(A)
Reliability
Lift = Confidence / Support(B)
Above 1 means positive association
Leverage = P(A and B) − P(A) P(B)
Difference from independence
Example
100 baskets: bread 40, butter 25, both 20 → support 0.20, confidence 0.50, lift 0.50 / 0.25 = 2 — strong positive rule.
- 1
Find frequent 1-itemsets
- 2
Join to form candidate k-itemsets
- 3
Prune candidates with an infrequent subset
- 4
Count support; keep frequent
- 5
Generate rules meeting minimum confidence
- 6
Evaluate candidate rules with lift and domain sense
Topic 5
Finding similarity
Jaccard = size(A ∩ B) / size(A ∪ B)
Sets
Cosine = A·B / (norm(A) norm(B))
Vectors, text
Euclidean distance
Numeric
- MinHash and locality-sensitive hashing (LSH) find similar items quickly in huge collections.
Topic 6
Recommendation systems
Collaborative filtering
Users who agreed before will agree again (user–user, item–item, matrix factorisation)
Cold start for new users and items
Content-based
Recommend items similar to those the user liked, using item features
Over-specialisation
Knowledge-based
Rules and constraints from explicit requirements (cars, houses)
Needs knowledge engineering
Hybrid
Combine approaches
More complex
Key terms
- WSS
- Within-cluster sum of squares
- AUC
- Area under the ROC curve
- Lift
- Confidence divided by the consequent's support
- Jaccard similarity
- Overlap of two sets
- Collaborative filtering
- Recommendations from similar users' behaviour
Quick revision
- K-means; elbow; silhouette.
- Trees: entropy, gain, Gini; confusion matrix, ROC.
- Naive Bayes.
- Support, confidence, lift; Apriori.
- Jaccard, cosine, LSH; collaborative, content, knowledge, hybrid recommenders.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.What is the elbow method?
- Q2.Define lift.
- Q3.What is the ROC curve?
- Q4.Define Jaccard similarity.
- Q5.What is the cold-start problem?
- Q6.What is a hybrid recommender?
Long-answer questions
- Q1.Explain k-means clustering and its diagnostics.
- Q2.Explain the Apriori algorithm and the evaluation of candidate rules.
- Q3.Explain the types of recommendation systems.
Stuck on this unit?
Message SBS on WhatsApp for help with Big Data Analytics, or to ask about studying M.Sc IT at Synetic.
