Unit 2 of 4 · M.Sc IT Sem 4

Unit 2: Advanced analytical methods

Big Data Analytics notes · PTU syllabus (PGCA1947)

3 min read6 topics9 exam questions
On this page
  1. Unit summary
  2. K-means clustering: use cases and diagnostics
  3. Decision trees and their evaluation
  4. Naive Bayes
  5. Association rules and Apriori
  6. Finding similarity
  7. Recommendation systems
  8. Key terms
  9. Quick revision
  10. Important questions

Unit summary

Analytical methods turn big data into insight. This unit covers k-means clustering use cases and diagnostics, decision trees and their evaluation, Naive Bayes, association rules and Apriori with rule evaluation, finding similarity, and recommendation systems.

After this unit you can

  • Apply k-means and diagnose cluster quality
  • Build and evaluate decision trees and Naive Bayes
  • Evaluate association rules with support, confidence and lift
  • Compare recommendation system approaches

PTU syllabus topics

  • K-means clustering use cases and diagnostics
  • decision tree algorithms and evaluation
  • Naive Bayes classification
  • association rules and the Apriori algorithm
  • evaluation of candidate rules
  • finding association and similarity
  • collaborative/content-based/knowledge-based/hybrid recommendation systems
ComparisonTypes of recommender systems
Based on
Example

Collaborative filtering

Similar users' behaviour

People who bought this also bought

Content-based

Item features

Similar movies by genre

Knowledge-based

Explicit rules and needs

Configure a laptop by requirements

Hybrid

A mix

Netflix-style recommendations

1

Topic 1

K-means clustering: use cases and diagnostics

ProcessK-means
  1. 1Choose k
  2. 2Initialise centroids
  3. 3Assign to nearest centroid
  4. 4Recompute means
  5. 5Repeat until stable
Key termsDiagnostics
WSS (within sum of squares)
Plot against k; elbow suggests k
Silhouette
−1 to 1; higher means well-separated
Scaling
Standardise features
Use cases
Customer segmentation, image compression, document grouping
2

Topic 2

Decision trees and their evaluation

Key formulasTree measures
  • Entropy = −Σ p log2 p

    Impurity

  • Gain = Entropy(parent) − weighted Entropy(children)

    Split choice

  • Gini = 1 − Σ p²

    CART impurity

Key termsAlgorithms and evaluation
ID3, C4.5, CART
Common algorithms
Confusion matrix
TP, FP, TN, FN
Accuracy, precision, recall
Performance
ROC curve and AUC
Trade-off across thresholds
Pruning
Avoid overfitting
3

Topic 3

Naive Bayes

Key formulasNaive Bayes
  • P(C given X) ∝ P(C) × Π P(xi given C)

    Independence assumption

  • Laplace smoothing

    Add 1 to counts

  • Fast, scales well to large text data (spam filtering, sentiment).
4

Topic 4

Association rules and Apriori

Key formulasRule measures
  • Support(A → B) = P(A and B)

    Frequency

  • Confidence = Support(A and B) / Support(A)

    Reliability

  • Lift = Confidence / Support(B)

    Above 1 means positive association

  • Leverage = P(A and B) − P(A) P(B)

    Difference from independence

Example

100 baskets: bread 40, butter 25, both 20 → support 0.20, confidence 0.50, lift 0.50 / 0.25 = 2 — strong positive rule.

ProcessApriori
  1. 1

    Find frequent 1-itemsets

  2. 2

    Join to form candidate k-itemsets

  3. 3

    Prune candidates with an infrequent subset

  4. 4

    Count support; keep frequent

  5. 5

    Generate rules meeting minimum confidence

  6. 6

    Evaluate candidate rules with lift and domain sense

5

Topic 5

Finding similarity

Key formulasSimilarity
  • Jaccard = size(A ∩ B) / size(A ∪ B)

    Sets

  • Cosine = A·B / (norm(A) norm(B))

    Vectors, text

  • Euclidean distance

    Numeric

  • MinHash and locality-sensitive hashing (LSH) find similar items quickly in huge collections.
6

Topic 6

Recommendation systems

ComparisonRecommenders
Idea
Limitation

Collaborative filtering

Users who agreed before will agree again (user–user, item–item, matrix factorisation)

Cold start for new users and items

Content-based

Recommend items similar to those the user liked, using item features

Over-specialisation

Knowledge-based

Rules and constraints from explicit requirements (cars, houses)

Needs knowledge engineering

Hybrid

Combine approaches

More complex

Key terms

WSS
Within-cluster sum of squares
AUC
Area under the ROC curve
Lift
Confidence divided by the consequent's support
Jaccard similarity
Overlap of two sets
Collaborative filtering
Recommendations from similar users' behaviour

Quick revision

  • K-means; elbow; silhouette.
  • Trees: entropy, gain, Gini; confusion matrix, ROC.
  • Naive Bayes.
  • Support, confidence, lift; Apriori.
  • Jaccard, cosine, LSH; collaborative, content, knowledge, hybrid recommenders.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.What is the elbow method?
  2. Q2.Define lift.
  3. Q3.What is the ROC curve?
  4. Q4.Define Jaccard similarity.
  5. Q5.What is the cold-start problem?
  6. Q6.What is a hybrid recommender?

Long-answer questions

  1. Q1.Explain k-means clustering and its diagnostics.
  2. Q2.Explain the Apriori algorithm and the evaluation of candidate rules.
  3. Q3.Explain the types of recommendation systems.

Stuck on this unit?

Message SBS on WhatsApp for help with Big Data Analytics, or to ask about studying M.Sc IT at Synetic.

WhatsApp us