Level I · Quantitative

Learning Module 10
Simple Linear Regression

Key Outcomes Summary & Practice Problems

Learning Outcomes

What you must be able to do

Curriculum Year: 2026

LOS 1

Describe a simple linear regression model, how OLS estimates regression coefficients, and interpret those coefficients.

LOS 2

Explain the four assumptions of SLR and how residuals and residual plots indicate if assumptions are violated.

LOS 3

Calculate and interpret R², F-statistic, and standard error; formulate and evaluate tests of fit and regression coefficients.

LOS 4

Describe ANOVA in regression analysis, interpret ANOVA results, and calculate the standard error of estimate (SEE).

LOS 5

Calculate and interpret the predicted value for the dependent variable, and a prediction interval, given an estimated linear regression model.

LOS 6

Describe different functional forms of simple linear regressions (log-lin, lin-log, log-log) and select the most appropriate form.

1 · The Simple Linear Regression Model & OLS

Simple linear regression (SLR) describes how one variable (Y) varies linearly with another (X), using the ordinary least squares (OLS) method — minimising the sum of squared residuals to fit the best line through the data.

SLR MODEL

Population model: Yᵢ = b₀ + b₁Xᵢ + εᵢ (i = 1, …, n)
Fitted line: Ŷᵢ = b̂₀ + b̂₁Xᵢ (estimated / predicted value)
Residual: eᵢ = Yᵢ − Ŷᵢ (observed minus predicted)

Y = dependent / explained variable (placed on vertical axis)
X = independent / explanatory variable (horizontal axis)
b₀ = intercept (population parameter)
b₁ = slope (population parameter)
εᵢ = error term (true underlying deviation from population line)

Key: error term ε ≠ residual eᵢ. Error = true population concept;
residual = sample-based deviation from the fitted line.

OLS FORMULAS

OLS criterion: minimize SSE = Σeᵢ² = Σ(Yᵢ − Ŷᵢ)²

Slope: b̂₁ = Σ(Yᵢ − Ȳ)(Xᵢ − X̄) / Σ(Xᵢ − X̄)² = Cov(Y,X) / Var(X)
Intercept: b̂₀ = Ȳ − b̂₁X̄

ROA/CAPEX example (n=6): X̄ = 6.1%, Ȳ = 12.5%
Σ(Yᵢ−Ȳ)(Xᵢ−X̄) = 153.30 ; Σ(Xᵢ−X̄)² = 122.64
b̂₁ = 153.30/122.64 = 1.25
b̂₀ = 12.5 − (1.25 × 6.1) = 4.875
Fitted model: Ŷᵢ = 4.875 + 1.25 × CAPEXᵢ

Coefficient interpretation: The intercept (b̂₀ = 4.875) is the predicted value of Y when X = 0 — ROA when CAPEX = 0%. The slope (b̂₁ = 1.25) is the change in Y for a one-unit increase in X — for every 1% increase in CAPEX, ROA increases by 1.25%. The sign of the slope equals the sign of the correlation (same covariance numerator; different denominator). Sum of residuals always = 0 by construction: E(ε) = 0.

    • Slope vs correlation: same sign and driven by the same covariance, but the slope denominator is Var(X) while correlation denominator is σ_X × σ_Y.

    • Cross-sectional vs time-series: cross-sectional = many entities at one time (i = 1,…,n); time-series = one entity across time periods (t = 1,…,T).

    • Dependent variable (Y) is also called the explained variable; independent variable (X) is also called the explanatory or predictor variable.

    • Indicator (dummy) variable: when X = 0 or 1, the intercept = mean of Y when X = 0, and the slope = difference in group means. Interpreted like a t-test of difference in means.

2 · Four Assumptions of SLR — LHIN

1 · LINEARITY

The true relationship between Y and X is linear. X must be non-stochastic (non-random).Violation signal:curved or systematic pattern in residual plot (U-shape or inverted-U).

2 · HOMOSKEDASTICITY

Variance of residuals E(εᵢ²) = σ²_ε is constant across all i. Violation (heteroskedasticity):residuals cluster into groups with different variances — e.g., two interest rate regimes (Regime 1 slope 1.02, Regime 2 slope −0.28).

3 · INDEPENDENCE

Observations (Y,X pairs) are uncorrelated; residuals are uncorrelated across observations.Violation (autocorrelation):seasonal or trending pattern in residuals — e.g., quarterly revenues with Q4 spikes.

4 · NORMALITY

Regression residuals are normally distributed (not the raw data).Violation impact:non-normality is most problematic for small samples. CLT relaxes this assumption for large n.

    • Residual plots are the diagnostic tool: always plot residuals against X (and against time for time-series). Random scatter = assumptions satisfied. Any pattern = potential violation.

    • Linearity violation in residuals: systematic curved pattern → model is wrong functional form → consider log transformations.

    • Heteroskedasticity in practice: when data span different central bank policy regimes, different market states, or different time periods, variance of residuals may differ substantially between regimes.

    • Autocorrelation indicator: seasonal spikes (e.g., Q4 revenues jump, then fall) in the residual plot suggest correlated residuals, violating independence.

    • Outliers and normality: a single outlier (e.g., a data entry error) can dramatically alter estimated coefficients, R², and standard errors. Always examine outliers carefully before including in regression.

3 · Goodness of Fit: R², F-statistic & Standard Error

SS DECOMPOSITION

SST = SSR + SSE (Total = Explained + Unexplained)

SST = Σ(Yᵢ − Ȳ)² = total variation in Y
SSR = Σ(Ŷᵢ − Ȳ)² = variation explained by X (regression)
SSE = Σ(Yᵢ − Ŷᵢ)² = unexplained variation (residuals)

ROA/CAPEX example: SST=239.50 = SSR(191.625) + SSE(47.875)

R² (COEFF OF DETERMINATION)

Proportion of Y's variation explained by X:

R² = SSR / SST (ranges 0–100%)

In SLR only: R² = r² (square of pairwise correlation)

ROA example: R² = 191.625/239.50 = 80.01%
r = 0.8945 → r² = 0.8001 ✓ (confirms R² = r² in SLR)

Descriptive, not a statistical test. Use F-test to test significance.

F-STATISTIC

Tests H₀: b₁ = 0 vs Hₐ: b₁ ≠ 0 (joint test of slope = 0)

MSR = SSR/1 = SSR (for SLR with k=1)
MSE = SSE/(n−2)
F = MSR / MSE with df (1, n−2); right-tailed only

In SLR: F = t² (F-statistic = square of slope t-statistic)

ROA: F = 191.625 / (47.875/4) = 191.625/11.969 = 16.01
Critical F (5%, df 1,4) = 7.71 → 16.01 > 7.71 → Reject H₀

Also: t = 4.001; t² = 16.01 = F ✓

T-TEST FOR SLOPE

t = (b̂₁ − B₁) / s_{b̂₁} with df = n − 2

Standard error of slope: s_{b̂₁} = sₑ / √Σ(Xᵢ−X̄)²

Standard error of estimate (SEE): sₑ = √MSE = √(SSE/(n−2))

To test intercept: t = (b̂₀ − B₀) / s_{b̂₀} with df = n − 2

ROA slope test: H₀:b₁=0, df=4, critical t=±2.776
sₑ = √11.969 = 3.4596 ; s_{b̂₁} = 3.4596/√122.64 = 0.3124
t = (1.25−0)/0.3124 = 4.001 > 2.776 → Reject H₀

ANOVA TABLE STRUCTURE

Source

Sum of Squares

Degrees of Freedom

Mean Square

F-Statistic

Regression

SSR = Σ(Ŷᵢ−Ȳ)²

1 (k)

MSR = SSR/1

MSR/MSE

Error (Residual)

SSE = Σ(Yᵢ−Ŷᵢ)²

n − 2 (n−k−1)

MSE = SSE/(n−2)

Total

SST = Σ(Yᵢ−Ȳ)²

n − 1

—

—

Variance of Y = SST/(n−1). SEE = sₑ = √MSE = √(SSE/(n−2)). The smaller the SEE, the better the fit.

4 · Prediction & Prediction Intervals

POINT PREDICTION

Given forecasted X_f, the predicted dependent variable is:
Ŷ_f = b̂₀ + b̂₁ × X_f

ROA example: CAPEX_f = 6.0%
Ŷ_f = 4.875 + (1.25 × 6.0) = 12.375%

PREDICTION INTERVAL

Standard error of forecast (wider than SE of mean estimate):

s_f = sₑ × √[1 + 1/n + (X_f − X̄)² / Σ(Xᵢ−X̄)²]

Prediction interval: Ŷ_f ± t_{α/2, n−2} × s_f

Three factors that widen s_f:
(1) Larger sₑ (poorer fit);
(2) Smaller n (fewer observations);
(3) X_f far from X̄ (forecasting far from the mean of X)

ROA: X_f=6, X̄=6.1, s_f=3.459588×√1.166748 = 3.737
95% PI: 12.375 ± 2.776(3.737) = {2.00, 22.75}

    • Prediction interval vs confidence interval for mean: the prediction interval for an individual observation is always wider than the confidence interval for the mean prediction — because individual outcomes vary around the mean.

    • Interval widens away from X̄: the minimum forecast standard error occurs when X_f = X̄ and increases as X_f moves farther from the sample mean. Never extrapolate far outside the range of observed X values.

5 · Functional Forms for Non-Linear Relationships

When the Y–X relationship is non-linear, we transform the variables using natural logarithms. The resulting models are still linear in the parameters (b₀ and b₁), so OLS still applies.

LOG-LIN MODEL
Log-Lin
ln(Yᵢ) = b₀ + b₁Xᵢ

Y in log form; X linear. Slope = relative change in Y per unit increase in X. Use when Y grows exponentially with X (e.g., revenues vs time at constant growth rate).

LIN-LOG MODEL
Lin-Log
Yᵢ = b₀ + b₁ ln(Xᵢ)

Y linear; X in log form. Slope = absolute change in Y per relative change in X. Use when Y has diminishing response to increasing X (e.g., profit margin vs unit sales).

LOG-LOG MODEL
Log-Log
ln(Yᵢ) = b₀ + b₁ ln(Xᵢ)

Both in log form. Slope = elasticity (relative % change in Y per % change in X). Use for economic elasticities (e.g., revenues vs advertising spend).

Model

When to Use

Slope Interpretation

Forecasting Note

Lin-Lin (plain SLR)

Linear Y–X relationship; residuals random

1 unit ↑ in X → b₁ units ↑ in Y

Ŷ = b̂₀ + b̂₁X_f

Log-Lin

Y grows exponentially with X

1 unit ↑ in X → b₁×100% change in Y

Ŷ = e^(b̂₀ + b̂₁X_f) (take antilog)

Lin-Log

Diminishing returns in X (concave relationship)

1% ↑ in X → b₁/100 units ↑ in Y

Ŷ = b̂₀ + b̂₁ ln(X_f)

Log-Log

Both Y and X grow proportionally; need elasticity

1% ↑ in X → b₁% ↑ in Y (elasticity)

Ŷ = e^(b̂₀ + b̂₁ ln(X_f)) (take antilog)

Choosing functional form: Compare models using: (1) higher R², (2) higher F-statistic, (3) lower SEE (sₑ), (4) random residuals. Important: can only directly compare R² and SEE between models with the same dependent variable (same units). Cannot compare a log-lin model with a lin-lin model directly using R² since the Y variables differ (ln Y vs Y).