Unit 3: Association rule mining and classification
Data Warehousing & Mining notes · PTU syllabus (BSIT504/BSBC501)
On this page
- Unit summary
- Market basket analysis
- Association rules and the Apriori algorithm
- Apriori worked example
- Mining multilevel association rules
- From association to correlation analysis
- Constraint-based association mining
- Classification by decision tree induction
- Attribute selection measures: worked example
- Key terms
- Quick revision
- Important questions
Unit summary
Data mining finds patterns no one thought to look for. This unit covers market basket analysis, the Apriori algorithm, mining multilevel association rules, moving from association to correlation analysis, constraint-based association mining, classification by decision tree induction and attribute selection measures.
After this unit you can
- Explain market basket analysis and association rule measures
- Apply the Apriori algorithm
- Explain multilevel, correlation and constraint-based association mining
- Build decision trees using attribute selection measures
PTU syllabus topics
- Market basket analysis
- Apriori algorithm
- mining multilevel association rules
- association-to-correlation analysis
- constraint-based association mining
- classification by decision tree
- attribute selection measures
Support
Transactions with A and B / all transactions
Confidence
Support(A and B) / Support(A)
Lift
Confidence(A ⇒ B) / Support(B)
Lift > 1 means positive association
Apriori property
Every subset of a frequent itemset is frequent
Topic 1
Market basket analysis
- Market basket analysis studies customer transactions to find items bought together — used for shelf layout, cross-selling, bundles and recommendations ("customers who bought this also bought").
Topic 2
Association rules and the Apriori algorithm
An association rule X ⇒ Y means customers who buy X tend to buy Y (market basket analysis).
Support
Transactions containing X and Y / total transactions
Confidence
Support(X ∪ Y) / Support(X)
Lift
Confidence / Support(Y)
Lift > 1 means positive association
- 1
Find frequent 1-itemsets
Support ≥ minimum
- 2
Join to form candidate k-itemsets
- 3
Prune
Drop candidates with an infrequent subset
- 4
Count support and keep frequent ones
- 5
Repeat until no new itemsets
- 6
Generate rules with confidence ≥ minimum
Example
In 5 transactions, {bread, butter} appears in 3 and bread in 4: support = 60%, confidence(bread ⇒ butter) = 3/4 = 75%.
Topic 3
Apriori worked example
| TID | Items |
|---|---|
| T1 | Bread, Milk |
| T2 | Bread, Butter, Eggs |
| T3 | Milk, Butter, Cola |
| T4 | Bread, Milk, Butter |
| T5 | Bread, Milk, Cola |
- Minimum support = 3 transactions (60%).
- 1-itemsets: Bread 4, Milk 4, Butter 3, Cola 2, Eggs 1 → frequent: Bread, Milk, Butter.
- 2-itemsets: {Bread, Milk} 3, {Bread, Butter} 2, {Milk, Butter} 2 → frequent: {Bread, Milk}.
- Rules: Bread ⇒ Milk: confidence 3/4 = 75%; Milk ⇒ Bread: 3/4 = 75%.
- Apriori property: every subset of a frequent itemset must be frequent — so any candidate with an infrequent subset is pruned. Limitations: many candidates and repeated database scans; FP-growth avoids candidate generation with an FP-tree.
Topic 4
Mining multilevel association rules
- Items form a concept hierarchy (food → bread → wheat bread). Rules at low levels may lack support, while high-level rules may be too general.
Uniform support
Same minimum support at every level
Simple, but misses low-level rules or produces too many high-level ones
Reduced support
Lower minimum support at lower levels
Finds meaningful detailed rules
- Redundancy filtering: remove a rule such as "milk ⇒ wheat bread" if it is expected from its ancestor "milk ⇒ bread".
Topic 5
From association to correlation analysis
- High support and confidence can still mislead: if 75% of all customers buy coffee anyway, "tea ⇒ coffee" with 70% confidence is actually a negative link.
Lift
lift(A, B) = P(A ∪ B) ÷ (P(A) × P(B)) = confidence(A ⇒ B) ÷ support(B)
Interpretation
Lift > 1 positive correlation; = 1 independent; < 1 negative
Chi-square
Compares observed and expected counts in a contingency table
Example
Of 10,000 transactions, 6,000 include games, 7,500 videos, 4,000 both. Lift = 0.40 ÷ (0.60 × 0.75) = 0.89 < 1 — games and videos are negatively correlated despite 66% confidence.
Topic 6
Constraint-based association mining
- Knowledge-type
- Mine association, correlation or classification rules
- Data constraints
- Task-relevant data — sales in Punjab in 2025
- Dimension and level constraints
- Use only product and region at category level
- Interestingness constraints
- Minimum support, confidence, lift
- Rule constraints
- Rule form or item conditions — sum of prices less than ₹500
- Constraints focus mining on what users want and, when anti-monotonic (if a set fails, all supersets fail) or monotonic, can be pushed into the algorithm to prune the search.
Topic 7
Classification by decision tree induction
Classification learns a model from labelled data to predict the class of new data. It has two steps: training (learning) and testing (classification). A decision tree splits data on attributes; leaves are class labels. Attribute selection measures choose the best split:
- Information gain (ID3): reduction in entropy, Entropy = −Σ pᵢ log₂ pᵢ
- Gain ratio (C4.5): information gain normalised by split information
- Gini index (CART): 1 − Σ pᵢ²
Exam tip
In an exam, compute the entropy of the whole data set first, then the expected entropy after each split, and pick the attribute with the highest gain.
Topic 8
Attribute selection measures: worked example
- Data: 14 records, 9 "buys = yes", 5 "no". Entropy(D) = −(9/14) log₂(9/14) − (5/14) log₂(5/14) = 0.940.
- Split on age: youth (2 yes, 3 no), middle-aged (4 yes, 0 no), senior (3 yes, 2 no). Expected entropy = (5/14)(0.971) + (4/14)(0) + (5/14)(0.971) = 0.694.
- Gain(age) = 0.940 − 0.694 = 0.246 — higher than income (0.029), student (0.151) and credit rating (0.048), so age becomes the root.
Information gain
ID3
Favours attributes with many values
Gain ratio
C4.5
Corrects that bias with split information
Gini index
CART (binary splits)
Favours equal-sized partitions
- Tree pruning (pre-pruning or post-pruning) removes branches that reflect noise, preventing overfitting.
Key terms
- Market basket analysis
- Finding items frequently bought together
- Frequent itemset
- Itemset meeting minimum support
- Lift
- Ratio measuring correlation between items
- Multilevel association
- Rules at different levels of a concept hierarchy
- Information gain
- Reduction in entropy from a split
Quick revision
- Support, confidence, lift; Apriori steps and property; FP-growth.
- Multilevel: uniform vs reduced support; redundancy.
- Correlation: lift, chi-square; misleading strong rules.
- Constraints: knowledge, data, dimension, interestingness, rule; anti-monotone pruning.
- Decision tree induction; information gain, gain ratio, Gini; pruning.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.Define support and confidence.
- Q2.State the Apriori property.
- Q3.Why is reduced support used in multilevel mining?
- Q4.What does lift less than 1 indicate?
- Q5.What is a rule constraint?
- Q6.Which measure does C4.5 use?
Long-answer questions
- Q1.Explain the Apriori algorithm with an example.
- Q2.Explain mining of multilevel association rules.
- Q3.Explain correlation analysis and constraint-based association mining.
- Q4.Explain decision tree induction and attribute selection measures with an example.
Stuck on this unit?
Message SBS on WhatsApp for help with Data Warehousing & Mining, or to ask about studying B.Sc IT at Synetic.
