Learning Module 10
Simple Linear Regression
Key Outcomes Summary & Practice Problems
What you must be able to do
Curriculum Year: 2026
Describe a simple linear regression model, how OLS estimates regression coefficients, and interpret those coefficients.
Explain the four assumptions of SLR and how residuals and residual plots indicate if assumptions are violated.
Calculate and interpret R², F-statistic, and standard error; formulate and evaluate tests of fit and regression coefficients.
Describe ANOVA in regression analysis, interpret ANOVA results, and calculate the standard error of estimate (SEE).
Calculate and interpret the predicted value for the dependent variable, and a prediction interval, given an estimated linear regression model.
Describe different functional forms of simple linear regressions (log-lin, lin-log, log-log) and select the most appropriate form.
1 · The Simple Linear Regression Model & OLS
Simple linear regression (SLR) describes how one variable (Y) varies linearly with another (X), using the ordinary least squares (OLS) method — minimising the sum of squared residuals to fit the best line through the data.
Population model: Yᵢ = b₀ + b₁Xᵢ + εᵢ (i = 1, …, n)
Fitted line: Ŷᵢ = b̂₀ + b̂₁Xᵢ (estimated / predicted value)
Residual: eᵢ = Yᵢ − Ŷᵢ (observed minus predicted)
Y = dependent / explained variable (placed on vertical axis)
X = independent / explanatory variable (horizontal axis)
b₀ = intercept (population parameter)
b₁ = slope (population parameter)
εᵢ = error term (true underlying deviation from population line)
Key: error term ε ≠ residual eᵢ. Error = true population concept;
residual = sample-based deviation from the fitted line.
OLS criterion: minimize SSE = Σeᵢ² = Σ(Yᵢ − Ŷᵢ)²
Slope: b̂₁ = Σ(Yᵢ − Ȳ)(Xᵢ − X̄) / Σ(Xᵢ − X̄)² = Cov(Y,X) / Var(X)
Intercept: b̂₀ = Ȳ − b̂₁X̄
ROA/CAPEX example (n=6): X̄ = 6.1%, Ȳ = 12.5%
Σ(Yᵢ−Ȳ)(Xᵢ−X̄) = 153.30 ; Σ(Xᵢ−X̄)² = 122.64
b̂₁ = 153.30/122.64 = 1.25
b̂₀ = 12.5 − (1.25 × 6.1) = 4.875
Fitted model: Ŷᵢ = 4.875 + 1.25 × CAPEXᵢ
Coefficient interpretation: The intercept (b̂₀ = 4.875) is the predicted value of Y when X = 0 — ROA when CAPEX = 0%. The slope (b̂₁ = 1.25) is the change in Y for a one-unit increase in X — for every 1% increase in CAPEX, ROA increases by 1.25%. The sign of the slope equals the sign of the correlation (same covariance numerator; different denominator). Sum of residuals always = 0 by construction: E(ε) = 0.
Slope vs correlation: same sign and driven by the same covariance, but the slope denominator is Var(X) while correlation denominator is σ_X × σ_Y.
Cross-sectional vs time-series: cross-sectional = many entities at one time (i = 1,…,n); time-series = one entity across time periods (t = 1,…,T).
Dependent variable (Y) is also called the explained variable; independent variable (X) is also called the explanatory or predictor variable.
Indicator (dummy) variable: when X = 0 or 1, the intercept = mean of Y when X = 0, and the slope = difference in group means. Interpreted like a t-test of difference in means.
2 · Four Assumptions of SLR — LHIN
The true relationship between Y and X is linear. X must be non-stochastic (non-random).Violation signal:curved or systematic pattern in residual plot (U-shape or inverted-U).
Variance of residuals E(εᵢ²) = σ²_ε is constant across all i. Violation (heteroskedasticity):residuals cluster into groups with different variances — e.g., two interest rate regimes (Regime 1 slope 1.02, Regime 2 slope −0.28).
Observations (Y,X pairs) are uncorrelated; residuals are uncorrelated across observations.Violation (autocorrelation):seasonal or trending pattern in residuals — e.g., quarterly revenues with Q4 spikes.
Regression residuals are normally distributed (not the raw data).Violation impact:non-normality is most problematic for small samples. CLT relaxes this assumption for large n.
Residual plots are the diagnostic tool: always plot residuals against X (and against time for time-series). Random scatter = assumptions satisfied. Any pattern = potential violation.
Linearity violation in residuals: systematic curved pattern → model is wrong functional form → consider log transformations.
Heteroskedasticity in practice: when data span different central bank policy regimes, different market states, or different time periods, variance of residuals may differ substantially between regimes.
Autocorrelation indicator: seasonal spikes (e.g., Q4 revenues jump, then fall) in the residual plot suggest correlated residuals, violating independence.
Outliers and normality: a single outlier (e.g., a data entry error) can dramatically alter estimated coefficients, R², and standard errors. Always examine outliers carefully before including in regression.
3 · Goodness of Fit: R², F-statistic & Standard Error
SST = SSR + SSE (Total = Explained + Unexplained)
SST = Σ(Yᵢ − Ȳ)² = total variation in Y
SSR = Σ(Ŷᵢ − Ȳ)² = variation explained by X (regression)
SSE = Σ(Yᵢ − Ŷᵢ)² = unexplained variation (residuals)
ROA/CAPEX example: SST=239.50 = SSR(191.625) + SSE(47.875)
Proportion of Y's variation explained by X:
R² = SSR / SST (ranges 0–100%)
In SLR only: R² = r² (square of pairwise correlation)
ROA example: R² = 191.625/239.50 = 80.01%
r = 0.8945 → r² = 0.8001 ✓ (confirms R² = r² in SLR)
Descriptive, not a statistical test. Use F-test to test significance.
Tests H₀: b₁ = 0 vs Hₐ: b₁ ≠ 0 (joint test of slope = 0)
MSR = SSR/1 = SSR (for SLR with k=1)
MSE = SSE/(n−2)
F = MSR / MSE with df (1, n−2); right-tailed only
In SLR: F = t² (F-statistic = square of slope t-statistic)
ROA: F = 191.625 / (47.875/4) = 191.625/11.969 = 16.01
Critical F (5%, df 1,4) = 7.71 → 16.01 > 7.71 → Reject H₀
Also: t = 4.001; t² = 16.01 = F ✓
t = (b̂₁ − B₁) / s_{b̂₁} with df = n − 2
Standard error of slope: s_{b̂₁} = sₑ / √Σ(Xᵢ−X̄)²
Standard error of estimate (SEE): sₑ = √MSE = √(SSE/(n−2))
To test intercept: t = (b̂₀ − B₀) / s_{b̂₀} with df = n − 2
ROA slope test: H₀:b₁=0, df=4, critical t=±2.776
sₑ = √11.969 = 3.4596 ; s_{b̂₁} = 3.4596/√122.64 = 0.3124
t = (1.25−0)/0.3124 = 4.001 > 2.776 → Reject H₀
ANOVA TABLE STRUCTURE
Source | Sum of Squares | Degrees of Freedom | Mean Square | F-Statistic |
|---|---|---|---|---|
Regression | SSR = Σ(Ŷᵢ−Ȳ)² | 1 (k) | MSR = SSR/1 | MSR/MSE |
Error (Residual) | SSE = Σ(Yᵢ−Ŷᵢ)² | n − 2 (n−k−1) | MSE = SSE/(n−2) | |
Total | SST = Σ(Yᵢ−Ȳ)² | n − 1 | — | — |
Variance of Y = SST/(n−1). SEE = sₑ = √MSE = √(SSE/(n−2)). The smaller the SEE, the better the fit.
4 · Prediction & Prediction Intervals
Given forecasted X_f, the predicted dependent variable is:
Ŷ_f = b̂₀ + b̂₁ × X_f
ROA example: CAPEX_f = 6.0%
Ŷ_f = 4.875 + (1.25 × 6.0) = 12.375%
Standard error of forecast (wider than SE of mean estimate):
s_f = sₑ × √[1 + 1/n + (X_f − X̄)² / Σ(Xᵢ−X̄)²]
Prediction interval: Ŷ_f ± t_{α/2, n−2} × s_f
Three factors that widen s_f:
(1) Larger sₑ (poorer fit);
(2) Smaller n (fewer observations);
(3) X_f far from X̄ (forecasting far from the mean of X)
ROA: X_f=6, X̄=6.1, s_f=3.459588×√1.166748 = 3.737
95% PI: 12.375 ± 2.776(3.737) = {2.00, 22.75}
Prediction interval vs confidence interval for mean: the prediction interval for an individual observation is always wider than the confidence interval for the mean prediction — because individual outcomes vary around the mean.
Interval widens away from X̄: the minimum forecast standard error occurs when X_f = X̄ and increases as X_f moves farther from the sample mean. Never extrapolate far outside the range of observed X values.
5 · Functional Forms for Non-Linear Relationships
When the Y–X relationship is non-linear, we transform the variables using natural logarithms. The resulting models are still linear in the parameters (b₀ and b₁), so OLS still applies.
Y in log form; X linear. Slope = relative change in Y per unit increase in X. Use when Y grows exponentially with X (e.g., revenues vs time at constant growth rate).
Y linear; X in log form. Slope = absolute change in Y per relative change in X. Use when Y has diminishing response to increasing X (e.g., profit margin vs unit sales).
Both in log form. Slope = elasticity (relative % change in Y per % change in X). Use for economic elasticities (e.g., revenues vs advertising spend).
Model | When to Use | Slope Interpretation | Forecasting Note |
|---|---|---|---|
Lin-Lin (plain SLR) | Linear Y–X relationship; residuals random | 1 unit ↑ in X → b₁ units ↑ in Y | Ŷ = b̂₀ + b̂₁X_f |
Log-Lin | Y grows exponentially with X | 1 unit ↑ in X → b₁×100% change in Y | Ŷ = e^(b̂₀ + b̂₁X_f) (take antilog) |
Lin-Log | Diminishing returns in X (concave relationship) | 1% ↑ in X → b₁/100 units ↑ in Y | Ŷ = b̂₀ + b̂₁ ln(X_f) |
Log-Log | Both Y and X grow proportionally; need elasticity | 1% ↑ in X → b₁% ↑ in Y (elasticity) | Ŷ = e^(b̂₀ + b̂₁ ln(X_f)) (take antilog) |
Choosing functional form: Compare models using: (1) higher R², (2) higher F-statistic, (3) lower SEE (sₑ), (4) random residuals. Important: can only directly compare R² and SEE between models with the same dependent variable (same units). Cannot compare a log-lin model with a lin-lin model directly using R² since the Y variables differ (ln Y vs Y).