← All STAT 413 applets

The Bias–Variance Tradeoff

We fix a flexible model — a degree-8 polynomial fit with ridge regularization — and use the regularization strength λ as the single knob for model complexity. A large λ shrinks the coefficients hard, giving a stiff fit that is the same no matter which data we see (low variance) but systematically off (high bias). A small λ lets the curve chase every wiggle, so it is nearly unbiased but swings wildly from sample to sample (high variance). To see this, we draw B = 40 bootstrap resamples of the data, refit at each λ, and watch how the fitted curves spread out.

The dashed green curve is a stand-in for the unknown true function f(x): it is just a lightly-regularized fit on the full dataset (smallest λ). This is only an approximation of the truth — we never know f exactly — but it lets us decompose error into bias and variance. The irreducible noise σ2 is estimated as the residual variance of that reference fit.

Resampled fits at the current λ (spread = variance, gap to dashed = bias)

bootstrap fits average fit reference f(x) data

Bias2, Variance and Total error vs λ

Bias2 Variance Total σ2 floor
Complexity knob
slide left for large λ (stiff, high bias) · right for small λ (flexible, high variance)