Learning Module 8
Hypothesis Testing
Key Outcomes Summary & Practice Problems
What you must be able to do
Curriculum Year: 2026
Explain hypothesis testing and its components, including statistical significance, Type I and Type II errors, and the power of a test.
Construct hypothesis tests and determine their statistical significance, the associated Type I and Type II errors, and power of the test given a significance level.
Compare and contrast parametric and nonparametric tests, and describe situations where each is the more appropriate type of test.
1 · The Hypothesis Testing Framework
Hypothesis testing provides an objective, probability-based method for deciding whether sample data support or contradict a claim about a population parameter. We use it because risk and uncertainty mean we can never achieve certainty — only make well-calibrated probability statements.
Formulate H₀ (null) and Hₐ (alternative). Must be mutually exclusive and collectively exhaustive. State as "hoped-for" in Hₐ.
Choose the correct statistic (t, χ², F) and its probability distribution based on what is being tested and data assumptions.
Choose α (level of significance) — typically 1%, 5%, or 10%. This is the probability of Type I error you are willing to accept.
Identify critical values from the distribution. Specify rejection region(s): "Reject H₀ if |t| > critical value" or "if p-value < α."
Collect data, compute the test statistic from the sample using the appropriate formula.
Compare statistic to critical value (or p-value to α). Reject H₀ or fail to reject H₀. State the conclusion in context.
Decision via critical value vs p-value (equivalent methods):
• Critical value: Reject H₀ if |calculated statistic| > critical value
• p-value: Reject H₀ if p-value < α (level of significance)
The p-value is the smallest level of significance at which H₀ can be rejected — the area in the probability distribution beyond the calculated test statistic.
2 · Type I & II Errors, Significance, and Power
Probability = α (level of significance)
Reject H₀ when H₀ is actually TRUE. Incorrectly finding evidence for something that is not there. Also called a "false positive." The significance level α is the maximum probability of Type I error we allow.
Probability = β
Fail to reject H₀ when H₀ is actually FALSE. Missing real evidence. Also called a "false negative." The power of the test = 1 − β = probability of correctly rejecting a false null.
Decision | H₀ True | H₀ False |
|---|---|---|
Fail to Reject H₀ | ✓ Correct (1 − α) | ✗ Type II Error (β) |
Reject H₀ | ✗ Type I Error (α) | ✓ Power (1 − β) |
Trade-off between errors: decreasing α (stricter standard) reduces Type I error probability but increases Type II error probability β. We reject H₀ less often, including some cases where H₀ is actually false.
Power = 1 − β: the probability of correctly rejecting a false null hypothesis. A more powerful test is better. Power increases with larger sample size, larger true difference from H₀, and larger α.
Only way to reduce both errors simultaneously: increase the sample size n. More data improves the precision of estimates and raises power without sacrificing significance.
Confidence level = 1 − α: if α = 5%, confidence level = 95%. The confidence interval approach is equivalent to the two-tailed hypothesis test — H₀ is rejected if the hypothesized value falls outside the (1−α)% confidence interval.
3 · Common Test Statistics in Finance
Test Objective | Test Statistic | Distribution | Degrees of Freedom |
|---|---|---|---|
Single mean (σ unknown) | t = (X̄ − μ₀) / (s/√n) | t-distributed | n − 1 |
Difference of two means (independent, equal variances) | t = (X̄₁ − X̄₂) / √(sp²/n₁ + sp²/n₂) | t-distributed | n₁ + n₂ − 2 |
Mean of differences (paired/dependent) | t = (d̄ − μd₀) / (sd/√n) | t-distributed | n − 1 |
Single variance | χ² = (n−1)s² / σ₀² | Chi-square | n − 1 |
Equality of two variances | F = s₁² / s₂² | F-distributed | n₁−1, n₂−1 |
Correlation coefficient | t = r√(n−2) / √(1−r²) | t-distributed | n − 2 |
Used when population σ is unknown (the most common case):
t = (X̄ − μ₀) / (s / √n) with df = n − 1
Example — Sendar Equity Fund (n=24, X̄=1.5%, s=3.6%, μ₀=1.1%):
H₀: μ = 1.1% vs Hₐ: μ ≠ 1.1%; α = 5%; Critical t = ±2.069
t = (1.5 − 1.1) / (3.6/√24) = 0.4 / 0.7348 = 0.544
|0.544| < 2.069 → Fail to reject H₀
Independent samples, normal populations, equal but unknown variances:
Pooled variance: sp² = [(n₁−1)s₁² + (n₂−1)s₂²] / (n₁+n₂−2)
Test statistic: t = (X̄₁ − X̄₂) / √(sp²/n₁ + sp²/n₂)
df = n₁ + n₂ − 2
Tests whether a single population variance equals a hypothesised value σ₀²:
χ² = (n − 1) × s² / σ₀² with df = n − 1
Example: n=24, s=3.6% (s²=12.96), σ₀²=16 (σ₀=4%):
χ² = 23 × 12.96 / 16 = 18.63
Left-tailed test: Reject H₀ if χ² < critical value
Tests equality of variances from two independent normal populations:
F = s₁² / s₂² with df₁ = n₁−1, df₂ = n₂−1
Convention: place larger variance in numerator → F ≥ 1
Two-tailed: reject if F < F_lower OR F > F_upper
One-tailed right: reject if F > F_critical
Example: s²_Before = 4.644, s²_After = 3.919
F = 4.644 / 3.919 = 1.185; Critical (one-tail) = 1.17502 → Reject H₀
Paired comparisons test vs. independent samples: When the same entities (e.g., same companies, same days) are measured under two conditions, use the paired comparisons test (test of mean differences d̄). This is more powerful than the independent samples test because it removes variation due to the shared element. When samples are truly independent (e.g., two different groups of companies), use the pooled t-test for difference in means.
4 · Parametric vs. Nonparametric Tests
Parametric tests concern population parameters and rely on specific distributional assumptions (typically normality). Nonparametric tests make minimal distributional assumptions and can be used for ranks, ordinal data, or questions not about parameters.
Test Objective | Parametric Test | Nonparametric Alternative |
|---|---|---|
Single mean | t-test (or z-test) | Wilcoxon signed-rank test |
Difference in means (independent) | Pooled t-test | Mann–Whitney U test (Wilcoxon rank sum test) |
Mean differences (paired) | Paired t-test | Wilcoxon signed-rank test; Sign test |
Distribution of data | N/A | Kolmogorov-Smirnov test |
Randomness of data | N/A | Runs test |
Use nonparametric tests when:
Data violate distributional assumptions: small sample from a markedly non-normal population where t- or z-tests are inappropriate.
Outliers are present: extreme values distort parametric statistics (mean, variance) but not nonparametric ones (based on ranks or signs).
Data are ranked (ordinal): investment manager rankings, credit ratings, survey scales. Parametric tests require stronger measurement scales than ranks.
The hypothesis is not about a parameter: testing for randomness (runs test), testing whether a sample came from a particular distribution.
Preference for parametric when assumptions are met: if distributional assumptions hold, parametric tests generally have more power (higher ability to reject a false null) than nonparametric equivalents.