Bias-Variance Decomposition: OLS, Ridge, Lasso

Bias-variance decomposition when fitting on the mtcars dataset. Irreducible noise is estimated directly from data, then, bias and variance are estimated by Monte Carlo and the irreducible noise ( simulated training sets).

\( \definecolor{biasc}{RGB}{216,27,96}\definecolor{varc}{RGB}{25,118,210}\definecolor{irrc}{RGB}{117,117,117} \mathbb{E}\big[(y-\hat f(x))^2\big] \;=\; \textcolor{biasc}{\mathrm{Bias}^2\!\big(\hat f(x)\big)} \;+\; \textcolor{varc}{\mathrm{Var}\!\big(\hat f(x)\big)} \;+\; \textcolor{irrc}{\sigma^2} \qquad \text{(at a new observation } y = f(x)+\varepsilon \text{)} \)
Method
\(\lambda = 0\) is exactly OLS (left edge of every curve)
Penalty
— uncheck to hide a curve (the y-axis rescales to what is visible)

Training set

\(\lambda\) squared error

Test set

\(\lambda\) squared error