Unit 3 of 4 · B.Sc IT Sem 5

Unit 3: Association rule mining and classification

Data Warehousing & Mining notes · PTU syllabus (BSIT504/BSBC501)

3 min read8 topics10 exam questions
On this page
  1. Unit summary
  2. Market basket analysis
  3. Association rules and the Apriori algorithm
  4. Apriori worked example
  5. Mining multilevel association rules
  6. From association to correlation analysis
  7. Constraint-based association mining
  8. Classification by decision tree induction
  9. Attribute selection measures: worked example
  10. Key terms
  11. Quick revision
  12. Important questions

Unit summary

Data mining finds patterns no one thought to look for. This unit covers market basket analysis, the Apriori algorithm, mining multilevel association rules, moving from association to correlation analysis, constraint-based association mining, classification by decision tree induction and attribute selection measures.

After this unit you can

  • Explain market basket analysis and association rule measures
  • Apply the Apriori algorithm
  • Explain multilevel, correlation and constraint-based association mining
  • Build decision trees using attribute selection measures

PTU syllabus topics

  • Market basket analysis
  • Apriori algorithm
  • mining multilevel association rules
  • association-to-correlation analysis
  • constraint-based association mining
  • classification by decision tree
  • attribute selection measures
Key formulasAssociation rule measures
  • Support

    Transactions with A and B / all transactions

  • Confidence

    Support(A and B) / Support(A)

  • Lift

    Confidence(A ⇒ B) / Support(B)

    Lift > 1 means positive association

  • Apriori property

    Every subset of a frequent itemset is frequent

1

Topic 1

Market basket analysis

  • Market basket analysis studies customer transactions to find items bought together — used for shelf layout, cross-selling, bundles and recommendations ("customers who bought this also bought").
2

Topic 2

Association rules and the Apriori algorithm

An association rule X ⇒ Y means customers who buy X tend to buy Y (market basket analysis).

Key formulasRule measures
  • Support

    Transactions containing X and Y / total transactions

  • Confidence

    Support(X ∪ Y) / Support(X)

  • Lift

    Confidence / Support(Y)

    Lift > 1 means positive association

ProcessApriori algorithm
  1. 1

    Find frequent 1-itemsets

    Support ≥ minimum

  2. 2

    Join to form candidate k-itemsets

  3. 3

    Prune

    Drop candidates with an infrequent subset

  4. 4

    Count support and keep frequent ones

  5. 5

    Repeat until no new itemsets

  6. 6

    Generate rules with confidence ≥ minimum

Example

In 5 transactions, {bread, butter} appears in 3 and bread in 4: support = 60%, confidence(bread ⇒ butter) = 3/4 = 75%.

3

Topic 3

Apriori worked example

TIDItems
T1Bread, Milk
T2Bread, Butter, Eggs
T3Milk, Butter, Cola
T4Bread, Milk, Butter
T5Bread, Milk, Cola
  • Minimum support = 3 transactions (60%).
  • 1-itemsets: Bread 4, Milk 4, Butter 3, Cola 2, Eggs 1 → frequent: Bread, Milk, Butter.
  • 2-itemsets: {Bread, Milk} 3, {Bread, Butter} 2, {Milk, Butter} 2 → frequent: {Bread, Milk}.
  • Rules: Bread ⇒ Milk: confidence 3/4 = 75%; Milk ⇒ Bread: 3/4 = 75%.
  • Apriori property: every subset of a frequent itemset must be frequent — so any candidate with an infrequent subset is pruned. Limitations: many candidates and repeated database scans; FP-growth avoids candidate generation with an FP-tree.
4

Topic 4

Mining multilevel association rules

  • Items form a concept hierarchy (food → bread → wheat bread). Rules at low levels may lack support, while high-level rules may be too general.
ComparisonSupport strategies
Approach
Effect

Uniform support

Same minimum support at every level

Simple, but misses low-level rules or produces too many high-level ones

Reduced support

Lower minimum support at lower levels

Finds meaningful detailed rules

  • Redundancy filtering: remove a rule such as "milk ⇒ wheat bread" if it is expected from its ancestor "milk ⇒ bread".
5

Topic 5

From association to correlation analysis

  • High support and confidence can still mislead: if 75% of all customers buy coffee anyway, "tea ⇒ coffee" with 70% confidence is actually a negative link.
Key formulasCorrelation measures
  • Lift

    lift(A, B) = P(A ∪ B) ÷ (P(A) × P(B)) = confidence(A ⇒ B) ÷ support(B)

  • Interpretation

    Lift > 1 positive correlation; = 1 independent; < 1 negative

  • Chi-square

    Compares observed and expected counts in a contingency table

Example

Of 10,000 transactions, 6,000 include games, 7,500 videos, 4,000 both. Lift = 0.40 ÷ (0.60 × 0.75) = 0.89 < 1 — games and videos are negatively correlated despite 66% confidence.

6

Topic 6

Constraint-based association mining

Key termsTypes of constraints
Knowledge-type
Mine association, correlation or classification rules
Data constraints
Task-relevant data — sales in Punjab in 2025
Dimension and level constraints
Use only product and region at category level
Interestingness constraints
Minimum support, confidence, lift
Rule constraints
Rule form or item conditions — sum of prices less than ₹500
  • Constraints focus mining on what users want and, when anti-monotonic (if a set fails, all supersets fail) or monotonic, can be pushed into the algorithm to prune the search.
7

Topic 7

Classification by decision tree induction

Classification learns a model from labelled data to predict the class of new data. It has two steps: training (learning) and testing (classification). A decision tree splits data on attributes; leaves are class labels. Attribute selection measures choose the best split:

  • Information gain (ID3): reduction in entropy, Entropy = −Σ pᵢ log₂ pᵢ
  • Gain ratio (C4.5): information gain normalised by split information
  • Gini index (CART): 1 − Σ pᵢ²

Exam tip

In an exam, compute the entropy of the whole data set first, then the expected entropy after each split, and pick the attribute with the highest gain.

8

Topic 8

Attribute selection measures: worked example

  • Data: 14 records, 9 "buys = yes", 5 "no". Entropy(D) = −(9/14) log₂(9/14) − (5/14) log₂(5/14) = 0.940.
  • Split on age: youth (2 yes, 3 no), middle-aged (4 yes, 0 no), senior (3 yes, 2 no). Expected entropy = (5/14)(0.971) + (4/14)(0) + (5/14)(0.971) = 0.694.
  • Gain(age) = 0.940 − 0.694 = 0.246 — higher than income (0.029), student (0.151) and credit rating (0.048), so age becomes the root.
ComparisonAttribute selection measures
Used in
Bias

Information gain

ID3

Favours attributes with many values

Gain ratio

C4.5

Corrects that bias with split information

Gini index

CART (binary splits)

Favours equal-sized partitions

  • Tree pruning (pre-pruning or post-pruning) removes branches that reflect noise, preventing overfitting.

Key terms

Market basket analysis
Finding items frequently bought together
Frequent itemset
Itemset meeting minimum support
Lift
Ratio measuring correlation between items
Multilevel association
Rules at different levels of a concept hierarchy
Information gain
Reduction in entropy from a split

Quick revision

  • Support, confidence, lift; Apriori steps and property; FP-growth.
  • Multilevel: uniform vs reduced support; redundancy.
  • Correlation: lift, chi-square; misleading strong rules.
  • Constraints: knowledge, data, dimension, interestingness, rule; anti-monotone pruning.
  • Decision tree induction; information gain, gain ratio, Gini; pruning.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.Define support and confidence.
  2. Q2.State the Apriori property.
  3. Q3.Why is reduced support used in multilevel mining?
  4. Q4.What does lift less than 1 indicate?
  5. Q5.What is a rule constraint?
  6. Q6.Which measure does C4.5 use?

Long-answer questions

  1. Q1.Explain the Apriori algorithm with an example.
  2. Q2.Explain mining of multilevel association rules.
  3. Q3.Explain correlation analysis and constraint-based association mining.
  4. Q4.Explain decision tree induction and attribute selection measures with an example.

Stuck on this unit?

Message SBS on WhatsApp for help with Data Warehousing & Mining, or to ask about studying B.Sc IT at Synetic.

WhatsApp us