Unit 3 of 4 · MBA Sem 3

Unit 3: Classification and association techniques

Data Mining for Business Decisions notes · PTU syllabus (MBA 941-18)

3 min read9 topics10 exam questions
On this page
  1. Unit summary
  2. Regression and correlation in data mining
  3. Decision trees
  4. Bayesian classification
  5. k-nearest neighbour and case-based reasoning
  6. Neural networks
  7. Support vector machines
  8. Market basket analysis and association rules
  9. Genetic algorithms and link analysis
  10. The fuzzy set approach
  11. Key terms
  12. Quick revision
  13. Important questions

Unit summary

Classification and association techniques turn data into decisions — who will buy, default or churn, and which products go together. This unit covers regression and correlation, decision trees, clustering, neural networks, market basket analysis and association rules, genetic algorithms and link analysis, support vector machines, Bayesian classification, k-nearest neighbour, case-based reasoning and the fuzzy set approach.

After this unit you can

  • Apply decision trees, Bayesian classification and k-NN
  • Explain neural networks, SVM, genetic algorithms and case-based reasoning
  • Apply market basket analysis and association rules
  • Explain link analysis and the fuzzy set approach

PTU syllabus topics

  • Regression and correlation
  • decision trees
  • clustering
  • neural networks
  • market basket analysis and association rules
  • genetic algorithms and link analysis
  • Support Vector Machine
  • Bayesian classification
  • k-Nearest Neighbour
  • case-based reasoning
  • fuzzy set approach
ComparisonClassification techniques
Idea
Strength

Decision tree

If-then splits

Easy to explain

Naive Bayes

Probabilities with Bayes' theorem

Fast, works with little data

k-NN

Nearest neighbours vote

No training phase

SVM

Maximum-margin boundary

Accurate in high dimensions

Neural network

Layers of neurons

Complex patterns

1

Topic 1

Regression and correlation in data mining

  • Correlation measures the strength of linear association; regression models a target as a function of predictors — used for numeric prediction and to identify drivers.
  • Logistic regression is a classification technique for binary targets (default or not).
2

Topic 2

Decision trees

  • A decision tree splits data on the attribute that best separates the classes; leaves give the predicted class.
Key formulasSplitting criteria
  • Entropy

    −Σ p log2 p

  • Information gain (ID3)

    Entropy before split − weighted entropy after

  • Gain ratio (C4.5)

    Information gain ÷ split information

  • Gini index (CART)

    1 − Σ p²

  • Pruning removes branches that fit noise, reducing overfitting.

Example

A node has 10 customers: 5 buy, 5 do not; entropy = 1. Splitting on "income high" gives two pure groups; information gain = 1 − 0 = 1, the best possible split.

3

Topic 3

Bayesian classification

Key formulasBayes' theorem
  • Posterior

    P(C given X) = P(X given C) × P(C) ÷ P(X)

  • Naive Bayes

    Assumes attributes are independent given the class: P(X given C) = product of P(xi given C)

  • Use: spam filtering, text classification, quick baseline models; fast and works well with many attributes even if independence is not exact.
4

Topic 4

k-nearest neighbour and case-based reasoning

  • k-NN: classify a new case by the majority class of its k nearest neighbours (Euclidean distance on standardised data); a "lazy learner" — no model is built in advance.
  • Case-based reasoning (CBR): solve new problems by retrieving similar past cases, reusing and adapting their solutions, and retaining the new case (retrieve, reuse, revise, retain) — used in help desks and diagnosis.
5

Topic 5

Neural networks

  • Artificial neural network: layers of interconnected nodes (input, hidden, output) with weights adjusted by backpropagation to minimise error.
  • Strengths: capture complex non-linear patterns — image, speech, fraud detection; weaknesses: need large data, black box, risk of overfitting. Deep learning uses many hidden layers.
6

Topic 6

Support vector machines

  • SVM finds the hyperplane that separates classes with the maximum margin; the closest points are the support vectors.
  • Kernel trick (polynomial, radial basis function) handles non-linear boundaries by mapping data to higher dimensions.
  • Use: text classification, credit scoring, image recognition; accurate but less interpretable.
7

Topic 7

Market basket analysis and association rules

Key formulasAssociation rule measures (rule A ⇒ B)
  • Support

    Transactions containing A and B ÷ total transactions

  • Confidence

    Transactions containing A and B ÷ transactions containing A

  • Lift

    Confidence ÷ support of B; above 1 means positive association

Example

In 1,000 bills, 200 contain bread, 150 contain butter and 100 contain both. Support = 10%, confidence (bread ⇒ butter) = 100 ÷ 200 = 50%, lift = 0.50 ÷ 0.15 = 3.33 — bread buyers are 3.3 times as likely to buy butter.

  • Apriori algorithm: every subset of a frequent itemset must be frequent — generate candidates level by level and prune. FP-growth avoids candidate generation.
  • Uses: store layout, cross-selling, bundles, recommendation ("frequently bought together").
8

Topic 8

Genetic algorithms and link analysis

  • Genetic algorithms: optimisation inspired by natural selection — a population of candidate solutions evolves through selection, crossover and mutation; used for scheduling, feature selection, portfolio optimisation.
  • Link analysis: examines relationships between entities in a network — PageRank, social network analysis, fraud rings in banking and insurance, telecom call graphs.
9

Topic 9

The fuzzy set approach

  • Fuzzy logic allows partial membership between 0 and 1 instead of crisp categories — a customer can be 0.7 "high income" and 0.3 "medium income".
  • Use: classification rules with linguistic terms ("if income is high and age is young then risk is low"), control systems, credit evaluation.

Key terms

Information gain
Reduction in entropy from a split
Naive Bayes
Bayesian classifier assuming independent attributes
Support vector
Data point closest to the separating hyperplane
Lift
Confidence of a rule divided by support of the consequent
Fuzzy set
Set allowing partial membership

Quick revision

  • Regression and logistic regression for prediction and classification.
  • Decision trees: entropy, information gain, Gini; pruning.
  • Naive Bayes; k-NN; CBR (retrieve, reuse, revise, retain).
  • Neural networks (backpropagation); SVM (max margin, kernels).
  • Association rules: support, confidence, lift; Apriori; genetic algorithms; link analysis; fuzzy sets.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.What is information gain?
  2. Q2.State Bayes' theorem.
  3. Q3.Why is k-NN called a lazy learner?
  4. Q4.What is the kernel trick in SVM?
  5. Q5.Define support, confidence and lift.
  6. Q6.What is a genetic algorithm?

Long-answer questions

  1. Q1.Explain decision tree induction with a suitable example.
  2. Q2.Explain Bayesian classification and the k-nearest neighbour method.
  3. Q3.Explain market basket analysis and the Apriori algorithm with an example.
  4. Q4.Discuss neural networks, SVM and genetic algorithms in data mining.

Stuck on this unit?

Message SBS on WhatsApp for help with Data Mining for Business Decisions, or to ask about studying MBA at Synetic.

WhatsApp us