Regularization Paths: OLS, Ridge & Lasso

Below is a real dataset: 32 cars road-tested by Motor Trend in 1974. Each colored line is one coefficient \(\hat{\beta}_j(\lambda)\). Drag the \(\lambda\) slider and watch: ridge only shrinks coefficients smoothly, while lasso sets them exactly to zero, one by one. OLS has no penalty, so \(\lambda\) does nothing to it.

OLS — \(\lambda\) does nothing

\(\lambda\)
\(\hat{\beta}_j\)

Ridge — shrinks, never exactly 0

\(\lambda\)
\(\hat{\beta}_j\)

Lasso — sparsity: exact zeros

\(\lambda\)
\(\hat{\beta}_j\)
Penalty strength (log scale)
Lasso
nonzero coefficients at this \(\lambda\):   OLS 10 / 10  ·  Ridge 10 / 10  ·  Lasso 10 / 10

Coefficients at the current \(\lambda\)  ■ OLS  ■ Ridge  ■ Lasso

\(\hat{\beta}_j\)
\( \text{OLS:} \quad \hat{\beta} = \arg\min_{\beta}\; \frac{1}{n}\sum_{i=1}^{n} \Big( y_i - \sum_{j=1}^{10} \beta_j z_{ij} \Big)^2 \)
\( \text{Ridge:} \quad \hat{\beta}(\lambda) = \arg\min_{\beta}\; \text{MSE}(\beta) + \lambda \sum_{j=1}^{10} \beta_j^2 \) — closed form; smooth shrinkage
\( \text{Lasso:} \quad \hat{\beta}(\lambda) = \arg\min_{\beta}\; \text{MSE}(\beta) + \lambda \sum_{j=1}^{10} |\beta_j| \) — coordinate descent (soft-thresholding); exact zeros, all 10 gone by \(\lambda \approx 10.3\)