Unit 3: Data mining fundamentals
Data Warehousing and Data Mining notes · PTU syllabus (PGCA1941)
On this page
Unit summary
Data mining discovers useful patterns in large data sets. This unit covers data mining functionalities, kinds of data mined, classification vs clustering, predictive vs descriptive mining, association rule mining, market basket analysis and the Apriori algorithm.
After this unit you can
- Describe data mining functionalities and the kinds of data mined
- Distinguish classification and clustering, predictive and descriptive mining
- Explain support, confidence and association rules
- Apply the Apriori algorithm
PTU syllabus topics
- Data mining functionalities
- mining different kinds of data
- classification vs clustering
- predictive vs descriptive mining
- association rule mining
- market basket analysis
- the Apriori algorithm
Support
Count(A and B) / total transactions
Confidence
Support(A and B) / Support(A)
Lift
Confidence(A ⇒ B) / Support(B)
Apriori principle
Subsets of frequent itemsets are frequent
Topic 1
Data mining and KDD
- 1
Data cleaning
- 2
Data integration
- 3
Data selection
- 4
Data transformation
- 5
Data mining
- 6
Pattern evaluation
- 7
Knowledge presentation
Topic 2
Data mining functionalities
- Characterisation and discrimination
- Summarise a class; compare classes
- Association
- Items occurring together
- Classification
- Assign records to predefined classes
- Prediction (regression)
- Predict numeric values
- Clustering
- Group similar records without labels
- Outlier analysis
- Find unusual records (fraud)
- Evolution analysis
- Trends over time
Topic 3
Mining different kinds of data
- Relational and warehouse data
- Tables and cubes
- Transactional
- Market baskets
- Time series and sequence
- Stock prices, click streams
- Spatial
- Maps, GPS
- Text and web
- Documents, links, logs
- Multimedia
- Images, audio, video
- Graph and social network
- Connections between entities
Topic 4
Classification vs clustering; predictive vs descriptive
Learning
Supervised — labelled training data
Unsupervised — no labels
Classes
Predefined
Discovered
Example
Spam or not spam
Customer segments
Goal
Predict unknown or future values
Describe patterns in existing data
Techniques
Classification, regression, time-series forecasting
Clustering, association, summarisation
Example
Will this customer churn?
Which products sell together?
Topic 5
Market basket analysis
- Market basket analysis studies customer transactions to find items bought together — used for shelf layout, cross-selling, bundles and recommendations ("customers who bought this also bought").
Topic 6
Association rules and the Apriori algorithm
An association rule X ⇒ Y means customers who buy X tend to buy Y (market basket analysis).
Support
Transactions containing X and Y / total transactions
Confidence
Support(X ∪ Y) / Support(X)
Lift
Confidence / Support(Y)
Lift > 1 means positive association
- 1
Find frequent 1-itemsets
Support ≥ minimum
- 2
Join to form candidate k-itemsets
- 3
Prune
Drop candidates with an infrequent subset
- 4
Count support and keep frequent ones
- 5
Repeat until no new itemsets
- 6
Generate rules with confidence ≥ minimum
Example
In 5 transactions, {bread, butter} appears in 3 and bread in 4: support = 60%, confidence(bread ⇒ butter) = 3/4 = 75%.
Topic 7
Apriori worked example
| TID | Items |
|---|---|
| T1 | Bread, Milk |
| T2 | Bread, Butter, Eggs |
| T3 | Milk, Butter, Cola |
| T4 | Bread, Milk, Butter |
| T5 | Bread, Milk, Cola |
- Minimum support = 3 transactions (60%).
- 1-itemsets: Bread 4, Milk 4, Butter 3, Cola 2, Eggs 1 → frequent: Bread, Milk, Butter.
- 2-itemsets: {Bread, Milk} 3, {Bread, Butter} 2, {Milk, Butter} 2 → frequent: {Bread, Milk}.
- Rules: Bread ⇒ Milk: confidence 3/4 = 75%; Milk ⇒ Bread: 3/4 = 75%.
- Apriori property: every subset of a frequent itemset must be frequent — so any candidate with an infrequent subset is pruned. Limitations: many candidates and repeated database scans; FP-growth avoids candidate generation with an FP-tree.
Key terms
- KDD
- Knowledge discovery in databases
- Support
- Fraction of transactions containing an itemset
- Confidence
- P(B given A) for rule A → B
- Frequent itemset
- Itemset meeting minimum support
- Apriori property
- Every subset of a frequent itemset is frequent
Quick revision
- KDD steps; functionalities; kinds of data.
- Classification vs clustering; predictive vs descriptive.
- Market basket, support, confidence, lift; Apriori join and prune.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.List the steps of KDD.
- Q2.Name four data mining functionalities.
- Q3.Distinguish classification and clustering.
- Q4.What is predictive mining?
- Q5.Define support and confidence.
- Q6.State the Apriori property.
Long-answer questions
- Q1.Explain data mining functionalities and the kinds of data mined.
- Q2.Distinguish predictive and descriptive mining, and classification and clustering.
- Q3.Explain the Apriori algorithm with an example.
Stuck on this unit?
Message SBS on WhatsApp for help with Data Warehousing and Data Mining, or to ask about studying M.Sc IT at Synetic.
