Unit 2: Hypothesis testing and descriptive analytics
Business Analytics for Decision Making notes · PTU syllabus (MBA 201-26)
On this page
- Unit summary
- Sampling distributions and standard error
- Hypothesis testing and its errors
- Parametric tests: Z, t, F and ANOVA
- Chi-square, goodness of fit and association of attributes
- Running the tests in SPSS
- Data visualisation techniques
- Introduction to clustering and PCA
- Key terms
- Quick revision
- Important questions
Unit summary
Decisions often rest on samples, so managers must judge whether differences are real or due to chance. This unit covers sampling distributions and standard error, errors in hypothesis testing, Z, t, F, chi-square, ANOVA and goodness-of-fit tests using SPSS, association of attributes, data visualisation techniques, and an introduction to clustering and PCA.
After this unit you can
- Explain sampling distributions, standard error and hypothesis testing errors
- Apply Z, t, F, chi-square and ANOVA tests and interpret SPSS output
- Measure the association of attributes
- Apply data visualisation, clustering and PCA
PTU syllabus topics
- Sampling distributions and standard error
- hypothesis testing errors
- Z-test/t-test/F-test/Chi-square test/ANOVA/goodness of fit using SPSS
- association of attributes
- data visualization techniques
- introduction to clustering and PCA
t-test
Difference in two means (small sample)
Do two branches differ in average sales?
ANOVA (F-test)
Difference in three or more means
Do four ad campaigns differ in response?
Chi-square
Association between categories
Is gender linked to product preference?
Z-test
Means or proportions, large samples
Has the defect rate changed?
Topic 1
Sampling distributions and standard error
- Sampling distribution: the probability distribution of a statistic (mean, proportion) over all possible samples of a given size.
Standard error of mean
σ ÷ √n
Standard error of proportion
√[p(1 − p) ÷ n]
Finite population correction
Multiply SE by √[(N − n) ÷ (N − 1)]
- Central Limit Theorem: for large samples (n ≥ 30), the sampling distribution of the mean is approximately normal with mean μ and standard error σ/√n, whatever the shape of the population.
Topic 2
Hypothesis testing and its errors
- 1
State null (H0) and alternative (H1) hypotheses
- 2
Choose level of significance (α — 5% or 1%)
- 3
Select the test statistic
Z, t, F, χ²
- 4
Determine the critical region
One-tailed or two-tailed
- 5
Compute the test statistic from sample data
- 6
Decide
Reject H0 if statistic falls in the critical region (or p-value < α)
- 7
Interpret in business terms
Meaning
Rejecting a true H0
Accepting a false H0
Example
Concluding a new drug works when it does not
Missing a drug that actually works
Controlled by
Level of significance
Sample size and power (1 − β)
Topic 3
Parametric tests: Z, t, F and ANOVA
| Test | Used for | Statistic |
|---|---|---|
| Z-test | Mean or proportion, large samples (n ≥ 30) or σ known | Z = (x̄ − μ) ÷ (σ/√n) |
| t-test (one sample) | Mean, small sample, σ unknown | t = (x̄ − μ) ÷ (s/√n), df = n − 1 |
| t-test (two independent samples) | Difference of two means | Pooled variance, df = n1 + n2 − 2 |
| Paired t-test | Before–after on the same units | t = d̄ ÷ (sd/√n) |
| F-test | Equality of two variances | F = s1² ÷ s2² (larger on top) |
| ANOVA | Equality of three or more means | F = MSB ÷ MSW |
Example
A sample of 25 bulbs has mean life 1,520 hours, s = 100; claim μ = 1,500. t = 20 ÷ 20 = 1.0 < 2.064 (5%, df 24) → do not reject the claim.
- 1Total variation (SST)
- 2Between-group variation (SSB)
- 3Within-group variation (SSW)
- 4Mean squares
MSB = SSB ÷ (k − 1); MSW = SSW ÷ (N − k)
- 5F = MSB ÷ MSW compared with table F
- Two-way ANOVA tests two factors (e.g., salesperson and region) and their interaction.
Topic 4
Chi-square, goodness of fit and association of attributes
- Chi-square (χ²) = Σ (O − E)² ÷ E — non-parametric.
- Uses: goodness of fit (does data follow a distribution? df = k − 1 − parameters estimated), test of independence in contingency tables (df = (r − 1)(c − 1)), test of population variance.
Example
Survey of 200: preference for brand X by gender. If the computed χ² = 7.2 with df = 1 and the 5% table value is 3.84, gender and brand preference are associated.
- Association of attributes: Yule's coefficient of association Q = (AD − BC) ÷ (AD + BC) — ranges from −1 to +1; coefficient of colligation; contingency coefficient.
Topic 5
Running the tests in SPSS
| Test | SPSS menu | Read in output |
|---|---|---|
| One-sample t | Analyze → Compare Means → One-Sample T Test | t, df, Sig. (2-tailed) |
| Independent-samples t | Analyze → Compare Means → Independent-Samples T Test | Levene's test, then t and Sig. |
| Paired t | Analyze → Compare Means → Paired-Samples T Test | Mean difference, t, Sig. |
| One-way ANOVA | Analyze → Compare Means → One-Way ANOVA | F and Sig.; post hoc (Tukey) |
| Chi-square association | Analyze → Descriptive Statistics → Crosstabs → Statistics → Chi-square | Pearson chi-square, Sig.; Phi, Cramer's V |
| Goodness of fit | Analyze → Nonparametric Tests → Legacy Dialogs → Chi-square | Chi-square, Sig. |
- Decision rule: if Sig. (p-value) is below α (0.05), reject H0.
- Z-test: for large samples the t-test result is practically identical; recent SPSS versions also offer one-sample and independent-samples proportions tests.
Example
An independent-samples t-test on satisfaction scores of two branches gives t = 2.41, Sig. = 0.018. Since 0.018 is below 0.05, the difference in mean satisfaction is statistically significant.
Topic 6
Data visualisation techniques
Comparison
Bar or column chart
Trend over time
Line chart
Composition
Stacked bar; pie only for a few parts
Distribution
Histogram, box plot
Relationship
Scatter plot, bubble chart, heat map
Geography
Filled or symbol maps
- Principles: start bar axes at zero, label clearly, minimise clutter, use colour for meaning, show the source.
- Tools: SPSS Chart Builder, Excel, Tableau, Power BI.
Topic 7
Introduction to clustering and PCA
- Cluster analysis: groups cases so that cases within a cluster are similar and clusters differ — used for customer segmentation.
Approach
Merges cases step by step; dendrogram
Assigns cases to k chosen centres, iteratively
Number of clusters
Decided from the dendrogram
Must be fixed in advance
Data size
Small to medium
Large
SPSS
Analyze → Classify → Hierarchical Cluster
Analyze → Classify → K-Means Cluster
- PCA: converts correlated variables into uncorrelated principal components ordered by variance explained; the first few often capture most information.
Example
A bank clusters 50,000 customers on balance, transactions and product holding into four segments — young digital users, salaried savers, business accounts and senior depositors — and designs offers for each.
Key terms
- Standard error
- Standard deviation of a sampling distribution
- p-value
- Probability of a result at least as extreme if H0 is true
- ANOVA
- Test comparing means of three or more groups
- Goodness of fit
- Test of whether data follow an expected distribution
- K-means
- Clustering that assigns cases to k centres
Quick revision
- SE of mean σ/√n; CLT.
- Type I (α) and Type II (β) errors.
- Z (large n), t (small n), F (variances), ANOVA (3+ means), χ² (association, goodness of fit).
- SPSS menus and decision rule (Sig. below 0.05 → reject H0).
- Charts by purpose; hierarchical vs k-means clustering; PCA.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.What is standard error?
- Q2.Distinguish Type I and Type II errors.
- Q3.When is a t-test used instead of a Z-test?
- Q4.What does the Sig. value in SPSS output mean?
- Q5.Which chart suits a trend over time?
- Q6.Distinguish hierarchical and k-means clustering.
Long-answer questions
- Q1.Explain the procedure of hypothesis testing and the errors involved.
- Q2.Explain the t-test and ANOVA with their SPSS procedures.
- Q3.Explain the chi-square test for goodness of fit and association of attributes.
- Q4.Discuss data visualisation techniques and an introduction to clustering and PCA.
Stuck on this unit?
Message SBS on WhatsApp for help with Business Analytics for Decision Making, or to ask about studying MBA at Synetic.
