Learning Module 9
Tests of Independence
Key Outcomes Summary & Practice Problems
What you must be able to do
Curriculum Year: 2026
Explain parametric and nonparametric tests of the hypothesis that the population correlation coefficient equals zero, and determine whether the hypothesis is rejected at a given level of significance.
Explain tests of independence based on contingency table data using the chi-square test statistic, and determine whether the hypothesis of independence is rejected.
1 Β· Parametric Test of Correlation: Pearson (t-test)
The Pearson correlation coefficient (also called bivariate or pairwise correlation) measures the strength of the linearrelationship between two continuous variables. When both variables are normally distributed, we test whether the population correlation Ο equals zero using a t-statistic.
Sample correlation between X and Y:
r_XY = s_XY / (s_X Γ s_Y)
where s_XY = sample covariance, s_X and s_Y = standard deviations
The covariance drives the sign of r:
Positive s_XY β r > 0 (variables move together)
Negative s_XY β r < 0 (variables move oppositely)
Test statistic: t = rβ(n β 2) / β(1 β rΒ²)
Distribution: t with n β 2 degrees of freedom
Hypotheses (most common β test for any relationship):
Two-sided: Hβ: Ο = 0 vs Hβ: Ο β 0
One-sided+: Hβ: Ο β€ 0 vs Hβ: Ο > 0
One-sidedβ: Hβ: Ο β₯ 0 vs Hβ: Ο < 0
Decision: Reject Hβ if |t_calc| > t_critical (two-tailed)
or if t_calc > t_critical (one-tailed right)
Worked example: n=33, r=0.43051, df=31, critical t=2.45282 (1%)
t = 0.43051β31 / β(1β0.18534) = 2.656 > 2.453 β Reject Hβ
Key insight β sample size and significance: The magnitude of r needed to reject Hβ decreases as n increases because (1) degrees of freedom increase β smaller critical values, and (2) the numerator rβ(nβ2) grows with n β larger test statistic. Example: r = 0.35 with n=12 gives t=1.182 (not significant at 5%), but the same r with n=32 gives t=2.046 (just significant). As n β β, even very small correlations become statistically significant β always consider economic significance alongside statistical significance.
Pearson is parametric: assumes both variables are normally distributed. If this assumption is violated, use the nonparametric Spearman rank correlation instead.
Degrees of freedom = n β 2: we lose 2 degrees of freedom because the test estimates two parameters (means of X and Y) from the data before computing the correlation.
Sign vs magnitude: the sign of r tells direction (positive = same direction, negative = opposite). The magnitude (closer to Β±1) tells strength. |r| = 0.35 is the same strength whether positive or negative.
Correlation matrix: when testing multiple pairwise correlations simultaneously, calculate a separate t-statistic for each pair and compare each to the critical value. Reject Hβ individually for each pair that exceeds the critical value.
2 Β· Nonparametric Test of Correlation: Spearman Rank
When the population departs meaningfully from normality, or when data contain outliers, the Spearman rank correlation coefficient (r_S) is appropriate. It is computed on the ranks of the observations, not their actual values β making it robust to outliers and distributional violations.
Uses actual observed values. Assumes bivariate normality. Sensitive to outliers. Measures linear relationship. Appropriate for continuous data from normal populations.
Uses ranks of observations. No normality assumption. Robust to outliers. Measures monotonic relationship (including nonlinear). Appropriate when normality is violated, data are ordinal, or outliers are present.
1. Rank X observations from largest (rank 1) to smallest (rank n)
Repeat for Y. Ties: assign average of tied ranks (e.g., 3 and 4 tied β 3.5)
2. For each pair i: compute d_i = rank(X_i) β rank(Y_i) and d_iΒ²
3. Spearman correlation:
r_S = 1 β [6 Γ Ξ£d_iΒ²] / [n(nΒ² β 1)]
Example: n=35, Ξ£dΒ² = 2202
r_S = 1 β 6(2202)/[35(1225β1)] = 1 β 13212/42840 = 0.6916
For large samples (n > 30), use same t-statistic as Pearson:
t = r_S Γ β(n β 2) / β(1 β r_SΒ²) with df = n β 2
For small samples (n β€ 30): requires specialized critical value tables
Example: r_S = 0.6916, n=35, df=33, critical t = Β±2.0345
t = 0.6916β33 / β(1β0.4783) = 3.974/0.7224 = 5.500 β Reject Hβ
Hβ: r_S = 0 vs Hβ: r_S β 0 (most common β test for any monotonic relationship)
When to use Spearman vs. Pearson: use Spearman when (1) normality assumption is violated, (2) outliers are present and influential, (3) data are ordinal or ranked rather than continuous, or (4) you suspect a monotonic but non-linear relationship.
Monotonic vs linear: Pearson measures linear association only. Spearman captures monotonic association β any consistently increasing or decreasing relationship, whether linear or curved.
Ties in rankings: when two observations share a rank (tied values), assign each the average of the ranks they would have occupied. For example, if the 3rd and 4th largest values are tied, both get rank 3.5.
Spearman for financial data: financial returns, expense ratios, and other financial variables are often bounded and non-normally distributed, making Spearman the preferred correlation measure for many investment applications.
3 Β· Tests of Independence: Contingency Tables & Chi-Square
When data are categorical or discrete (e.g., investment type, credit rating, ESG category), correlation cannot be used. Instead, a contingency table (two-way table) organises the frequencies and a chi-square test tests whether the two classification dimensions are independent.
OBSERVED FREQUENCIES: 1,594 ETFS BY SIZE AND INVESTMENT TYPE
Investment Type | Small-Cap | Mid-Cap | Large-Cap | Total |
|---|---|---|---|---|
Value | 50 | 110 | 343 | 503 |
Growth | 42 | 122 | 202 | 366 |
Blend | 56 | 149 | 520 | 725 |
Total | 148 | 381 | 1,065 | 1,594 |
Under Hβ (independence), expected count for each cell:
E_ij = (Row total_i Γ Column total_j) / Overall total
Example β small-cap Value ETFs:
E = (503 Γ 148) / 1,594 = 74,444 / 1,594 = 46.703
If observed = expected in all cells β ΟΒ² = 0 (perfect independence).
Any deviation makes ΟΒ² > 0. Squared deviations β always positive.
Therefore: only one rejection region (right tail)
Sum over all m cells of the scaled squared deviations:
ΟΒ² = Ξ£_m (O_ij β E_ij)Β² / E_ij
Degrees of freedom: df = (r β 1)(c β 1)
where r = number of row categories, c = number of column categories
ETF example: r=3 investment types, c=3 size groups
df = (3β1)(3β1) = 4; critical ΟΒ² at 5% = 9.4877
Calculated ΟΒ² = 32.080 > 9.488 β Reject Hβ
Conclusion: ETF size and investment type are NOT independent
SIX-STEP PROCESS FOR CHI-SQUARE TEST OF INDEPENDENCE
Hβ: classifications are independent (no relationship). Hβ: classifications are not independent (relationship exists).
ΟΒ² = Ξ£(O_ij β E_ij)Β² / E_ij. Chi-square distributed with (rβ1)(cβ1) degrees of freedom.
Specify Ξ± (commonly 5%). This determines the critical value from the chi-square table.
Right-tailed only. Reject Hβ if ΟΒ²_calc > ΟΒ²_critical. One rejection region (no left tail).
Compute expected frequencies for all cells, then sum (OβE)Β²/E across all m = rΓc cells.
Compare ΟΒ²_calc to critical value. State conclusion about independence in context.
Also called Pearson residual β measures how far each cell deviates:
Standardized residual = (O_ij β E_ij) / βE_ij
Positive β more observations than expected under independence
Negative β fewer observations than expected under independence
ETF example findings:
Medium-cap Growth: residual = +3.69 (more than expected)
Large-cap Growth: residual = β2.72 (fewer than expected)
Mosaic charts visualise these residuals with colour coding across all cells.
Chi-square is always right-tailed: because deviations are squared, ΟΒ² β₯ 0 always. Perfect independence β ΟΒ² = 0. Departures from independence β larger positive ΟΒ². There is no left-tail rejection region.
Degrees of freedom = (rβ1)(cβ1): NOT rΓc. A 3Γ3 table has df=4, not 9. A 2Γ2 table has df=1. A 2Γ3 table has df=2. Common exam error is using rΓc instead of (rβ1)(cβ1).
Independence means multiplicative: if two classifications are truly independent, the expected cell frequency equals (row total Γ column total) / grand total. Any meaningful deviation indicates a relationship.
Use for categorical data only: use the chi-square test of independence when data are categorical or discrete (investment type, ESG rating, sector classification). Use Pearson or Spearman correlation for continuous data.
Contrast with chi-square for variance: the chi-square distribution is used for two entirely different purposes in this curriculum β (1) testing a single population variance (LM8) and (2) testing independence in contingency tables (LM9). Do not confuse them.