Unit 3: Classification and association techniques
Data Mining for Business Decisions notes · PTU syllabus (MBA 941-18)
On this page
- Unit summary
- Regression and correlation in data mining
- Decision trees
- Bayesian classification
- k-nearest neighbour and case-based reasoning
- Neural networks
- Support vector machines
- Market basket analysis and association rules
- Genetic algorithms and link analysis
- The fuzzy set approach
- Key terms
- Quick revision
- Important questions
Unit summary
Classification and association techniques turn data into decisions — who will buy, default or churn, and which products go together. This unit covers regression and correlation, decision trees, clustering, neural networks, market basket analysis and association rules, genetic algorithms and link analysis, support vector machines, Bayesian classification, k-nearest neighbour, case-based reasoning and the fuzzy set approach.
After this unit you can
- Apply decision trees, Bayesian classification and k-NN
- Explain neural networks, SVM, genetic algorithms and case-based reasoning
- Apply market basket analysis and association rules
- Explain link analysis and the fuzzy set approach
PTU syllabus topics
- Regression and correlation
- decision trees
- clustering
- neural networks
- market basket analysis and association rules
- genetic algorithms and link analysis
- Support Vector Machine
- Bayesian classification
- k-Nearest Neighbour
- case-based reasoning
- fuzzy set approach
Decision tree
If-then splits
Easy to explain
Naive Bayes
Probabilities with Bayes' theorem
Fast, works with little data
k-NN
Nearest neighbours vote
No training phase
SVM
Maximum-margin boundary
Accurate in high dimensions
Neural network
Layers of neurons
Complex patterns
Topic 1
Regression and correlation in data mining
- Correlation measures the strength of linear association; regression models a target as a function of predictors — used for numeric prediction and to identify drivers.
- Logistic regression is a classification technique for binary targets (default or not).
Topic 2
Decision trees
- A decision tree splits data on the attribute that best separates the classes; leaves give the predicted class.
Entropy
−Σ p log2 p
Information gain (ID3)
Entropy before split − weighted entropy after
Gain ratio (C4.5)
Information gain ÷ split information
Gini index (CART)
1 − Σ p²
- Pruning removes branches that fit noise, reducing overfitting.
Example
A node has 10 customers: 5 buy, 5 do not; entropy = 1. Splitting on "income high" gives two pure groups; information gain = 1 − 0 = 1, the best possible split.
Topic 3
Bayesian classification
Posterior
P(C given X) = P(X given C) × P(C) ÷ P(X)
Naive Bayes
Assumes attributes are independent given the class: P(X given C) = product of P(xi given C)
- Use: spam filtering, text classification, quick baseline models; fast and works well with many attributes even if independence is not exact.
Topic 4
k-nearest neighbour and case-based reasoning
- k-NN: classify a new case by the majority class of its k nearest neighbours (Euclidean distance on standardised data); a "lazy learner" — no model is built in advance.
- Case-based reasoning (CBR): solve new problems by retrieving similar past cases, reusing and adapting their solutions, and retaining the new case (retrieve, reuse, revise, retain) — used in help desks and diagnosis.
Topic 5
Neural networks
- Artificial neural network: layers of interconnected nodes (input, hidden, output) with weights adjusted by backpropagation to minimise error.
- Strengths: capture complex non-linear patterns — image, speech, fraud detection; weaknesses: need large data, black box, risk of overfitting. Deep learning uses many hidden layers.
Topic 6
Support vector machines
- SVM finds the hyperplane that separates classes with the maximum margin; the closest points are the support vectors.
- Kernel trick (polynomial, radial basis function) handles non-linear boundaries by mapping data to higher dimensions.
- Use: text classification, credit scoring, image recognition; accurate but less interpretable.
Topic 7
Market basket analysis and association rules
Support
Transactions containing A and B ÷ total transactions
Confidence
Transactions containing A and B ÷ transactions containing A
Lift
Confidence ÷ support of B; above 1 means positive association
Example
In 1,000 bills, 200 contain bread, 150 contain butter and 100 contain both. Support = 10%, confidence (bread ⇒ butter) = 100 ÷ 200 = 50%, lift = 0.50 ÷ 0.15 = 3.33 — bread buyers are 3.3 times as likely to buy butter.
- Apriori algorithm: every subset of a frequent itemset must be frequent — generate candidates level by level and prune. FP-growth avoids candidate generation.
- Uses: store layout, cross-selling, bundles, recommendation ("frequently bought together").
Topic 8
Genetic algorithms and link analysis
- Genetic algorithms: optimisation inspired by natural selection — a population of candidate solutions evolves through selection, crossover and mutation; used for scheduling, feature selection, portfolio optimisation.
- Link analysis: examines relationships between entities in a network — PageRank, social network analysis, fraud rings in banking and insurance, telecom call graphs.
Topic 9
The fuzzy set approach
- Fuzzy logic allows partial membership between 0 and 1 instead of crisp categories — a customer can be 0.7 "high income" and 0.3 "medium income".
- Use: classification rules with linguistic terms ("if income is high and age is young then risk is low"), control systems, credit evaluation.
Key terms
- Information gain
- Reduction in entropy from a split
- Naive Bayes
- Bayesian classifier assuming independent attributes
- Support vector
- Data point closest to the separating hyperplane
- Lift
- Confidence of a rule divided by support of the consequent
- Fuzzy set
- Set allowing partial membership
Quick revision
- Regression and logistic regression for prediction and classification.
- Decision trees: entropy, information gain, Gini; pruning.
- Naive Bayes; k-NN; CBR (retrieve, reuse, revise, retain).
- Neural networks (backpropagation); SVM (max margin, kernels).
- Association rules: support, confidence, lift; Apriori; genetic algorithms; link analysis; fuzzy sets.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.What is information gain?
- Q2.State Bayes' theorem.
- Q3.Why is k-NN called a lazy learner?
- Q4.What is the kernel trick in SVM?
- Q5.Define support, confidence and lift.
- Q6.What is a genetic algorithm?
Long-answer questions
- Q1.Explain decision tree induction with a suitable example.
- Q2.Explain Bayesian classification and the k-nearest neighbour method.
- Q3.Explain market basket analysis and the Apriori algorithm with an example.
- Q4.Discuss neural networks, SVM and genetic algorithms in data mining.
Stuck on this unit?
Message SBS on WhatsApp for help with Data Mining for Business Decisions, or to ask about studying MBA at Synetic.
