Level I · Quantitative

Learning Module 7
Estimation and Inference

Key Outcomes Summary & Practice Problems

Learning Outcomes

What you must be able to do

Curriculum Year: 2026

LOS 1

Compare and contrast simple random, stratified random, cluster, convenience, and judgmental sampling and their implications for sampling error in an investment problem.

LOS 2

Explain the central limit theorem and its importance for the distribution and standard error of the sample mean.

LOS 3

Describe the use of resampling (bootstrap, jackknife) to estimate the sampling distribution of a statistic.

1 · Sampling Methods

A population is all members of a specified group. A sample is a subset drawn to estimate a population parameter. Sampling error is the difference between an observed sample statistic and the true population parameter it estimates — an unavoidable cost of using subsets instead of the full population.

PROBABILITY

Simple Random

Every element has equal probability of selection. Best when population is homogeneous. The foundation of all other probability methods.

PROBABILITY

Systematic

Select every k-th member from the population list. Used when population cannot be fully enumerated. Approximately random in practice.

PROBABILITY

Stratified Random

Divide population into strata; draw proportional random samples from each. Lower variance than simple random. Bond indexing key application.

PROBABILITY

Cluster

Divide into mini-population clusters; randomly select entire clusters. Most time/cost-efficient but lowest accuracy among probability methods.

NON-PROBABILITY

Convenience

Select based on ease of access. Fast and cheap. Significant risk of non-representative sample. Used in pilot studies or preliminary research.

NON-PROBABILITY

Judgmental

Expert handpicks elements using professional judgment. Risk of researcher bias. Used when expertise is critical or time is constrained (e.g., auditing).

Method

Type

Key Advantage

Key Risk / Limitation

Investment Application

Simple Random

Probability

Unbiased; each element equally likely

May miss important subgroups

Homogeneous datasets

Systematic

Probability

Practical when list is long

Periodicity bias if pattern exists

Large sorted databases

Stratified Random

Probability

Lower variance; subgroups guaranteed

Requires prior knowledge of strata

Bond indexing; fund replication

Cluster

Probability

Most cost/time efficient

Lower accuracy than stratified

Geographic market surveys

Convenience

Non-probability

Fastest; lowest cost

High non-representative risk

Pilot studies; preliminary research

Judgmental

Non-probability

Leverages expert knowledge

Researcher bias may skew results

Audit selection; specialist coverage

Stratified random sampling — bond indexing example: A portfolio manager replicating the Bloomberg Barclays US Government/Credit Index divides bonds into 3 issuer types × 10 maturity intervals × 2 coupon levels = 60 strata (cells). At least one bond is drawn from each cell → minimum 60 issues. Stratified sampling ensures all major risk factors (duration, sector, credit, coupon) are represented, unlike simple random sampling which could miss entire maturity buckets.

Sampling from different distributions — the Sharpe ratio trap: A manager who followed a low-risk strategy in Year 1 (Sharpe = 0.22) and a high-risk strategy in Year 2 (Sharpe = 0.22) appears to consistently outperform the benchmark (0.21). But pooling the 8 quarters together yields a Sharpe of only 0.199 — apparently underperforming. The problem: Year 1 and Year 2 returns come from different distributions (different μ and σ²). Pooling violates the assumption of a single homogeneous population. A larger sample is not always better — it must be drawn from the same distribution.

    • Probability vs. non-probability: probability sampling gives every element an equal chance of selection; non-probability sampling does not. All else equal, probability sampling yields more accurate and reliable estimates.

    • Sampling distribution: the distribution of all possible values that a statistic can assume when computed from samples of the same size drawn from the same population. The sample mean is itself a random variable with a sampling distribution.

    • Stratified > simple random (variance): by ensuring proportional representation of key subgroups, stratified sampling produces estimates with smaller variance (higher precision) than simple random sampling of equal size.

    • Cluster sampling trade-off: clusters are designed to be mini-representations of the entire population. Only sampled clusters are included (unlike stratified, where all strata contribute). This makes cluster sampling the most cost-efficient but least accurate probability method.

2 · Central Limit Theorem & Standard Error

The Central Limit Theorem (CLT) is one of the most powerful results in statistics — it tells us that regardless of the underlying population distribution, the sampling distribution of the sample mean becomes approximately normal as the sample size grows. This enables inference about population parameters even when we don't know the population's true distribution.

SMALL N = 10 · WIDE, IRREGULAR
Wide spread · Not yet normal
LARGE N = 300 · NARROW, NORMAL
Tight bell curve · Approximately normal
CLT STATEMENT

Given ANY population with mean μ and finite variance σ², the sampling
distribution of the sample mean X̄ from samples of size n will be:

X̄ ~ approximately Normal(μ, σ²/n) when n is large

Three properties guaranteed by CLT:
1. X̄ is approximately normally distributed (regardless of population shape)
2. Mean of X̄ = μ (population mean — X̄ is an unbiased estimator)
3. Variance of X̄ = σ²/n (shrinks as n grows — more precision)

Rule of thumb: CLT applies well when n ≥ 30; may need n >> 30
for very non-normal populations.

STANDARD ERROR

Standard error of the sample mean = std dev of the sampling distribution:

When population σ is known: σ_X̄ = σ / √n

When population σ is unknown (typical in practice):
s_X̄ = s / √n where s² = Σ(Xᵢ − X̄)² / (n − 1)

Key: σ_X̄ decreases as n increases → more precision with larger samples.
To halve the standard error, quadruple the sample size (because √4 = 2).

Standard deviation ≠ Standard error:
SD → describes data dispersion around the sample mean
SE → describes sampling precision of the estimated parameter

Why CLT matters for investment analysis: Financial return distributions are often non-normal (fat tails, skewness). Without CLT, we couldn't make reliable inferences about population means from samples. CLT allows analysts to construct confidence intervals and test hypotheses about the true population mean using the normal distribution, regardless of what the underlying return distribution looks like — as long as n ≥ 30 (or more for very skewed distributions).

    • CLT requires finite variance: the population must have finite variance σ² for CLT to apply. Distributions with infinite variance (e.g., Cauchy distribution) are exceptions — their sample means do not converge to normality.

    • Population normality NOT required: this is the CLT's power. Even if individual returns follow a highly skewed distribution, the mean of a large sample from that distribution will be approximately normally distributed.

    • Standard error vs. standard deviation distinction: if you want to describe how spread out your data are, use standard deviation. If you want to know how accurately your sample mean estimates the true population mean, use standard error (= s/√n).

    • Sample size effect: variance of X̄ = σ²/n — as n doubles, variance of X̄ halves and standard error decreases by 1/√2. Diminishing returns: going from n=100 to n=400 halves the standard error, but going from n=400 to n=1,600 halves it again.

3 · Resampling: Bootstrap & Jackknife

Resampling methods allow statistical inference without relying on analytical formulas like z- or t-statistics. They are particularly valuable when no closed-form formula exists for the estimator of interest (e.g., the standard error of the median).

🔄 BOOTSTRAP

Repeatedly draw resamples of the same size as the original sample, with replacement, from the observed data. Each item can be drawn multiple times; others may not appear at all. Compute the statistic (e.g., mean) for each resample. Build the bootstrap sampling distribution. Works for any estimator — including the median — where no analytical SE formula exists. Results vary across runs (random draws).

✂️ JACKKNIFE

Leave out one observation at a time (without replacement). For sample of size n, produce n resamples each of size n−1. Primarily used to reduce estimator bias. Also used to find standard error and confidence intervals. Produces the same results every run(deterministic — no randomness). Requires exactly n repetitions for a sample of size n.

BOOTSTRAP SE FORMULA

Standard error of the sample mean via bootstrap (B resamples):

s_X̄ = √[ (1/(B−1)) × Σᵦ(θ̂ᵦ − θ̄)² ]

Where:
B = number of bootstrap resamples (e.g., 1,000)
θ̂ᵦ = mean of the b-th resample
θ̄ = mean of all B resample means

Worked example: B=1,000; θ̄ = −0.01367; Σ(θ̂ᵦ−θ̄)² = 1.94143
s_X̄ = √(1.94143 / 999) = √0.001943 = 0.04408

Dimension

Bootstrap

Jackknife

Sampling method

With replacement from original sample

Leave-one-out (without replacement)

Resample size

Same size n as original sample

n − 1 (one observation removed each time)

Number of repetitions

Analyst chooses B (commonly 1,000+)

Always exactly n repetitions

Results across runs

Random — different results each run

Deterministic — same results every run

Primary use

SE of any estimator; confidence intervals

Bias reduction; SE and confidence intervals

Analytical formula needed?

No — model-free / non-parametric

No — also non-parametric

Key advantage

Applicable to any complicated estimator (median, skewness, etc.)

Deterministic; well-suited for bias correction

    • Bootstrap for the median: the formula s_X̄ = s/√n applies only to the sample mean, not the median. To find the SE of the sample median, bootstrap resampling is the appropriate tool — no analytical formula is needed.

    • Non-parametric resampling: bootstrap (and jackknife) are often called "model-free" resampling because they do not assume any particular distribution for the data. The empirical distribution of the observed sample serves as the population proxy.

    • Finance applications of bootstrap: historical simulation for asset allocation (drawing historical return scenarios with replacement); gauging an investment strategy's performance against a benchmark; estimating risk metrics (VaR, CVaR) without distributional assumptions.

    • Both methods provide statistical estimates, not exact results — they are complements to analytical methods, not replacements. Analytical methods (when available) provide more insight into cause-and-effect relationships.