Below is a real dataset: 32 cars road-tested by Motor Trend in 1974.
Each colored line is one coefficient \(\hat{\beta}_j(\lambda)\). Drag the \(\lambda\) slider and watch: ridge only shrinks
coefficients smoothly, while lasso sets them exactly to zero, one by one.
OLS has no penalty, so \(\lambda\) does nothing to it.
\( \text{OLS:} \quad \hat{\beta} = \arg\min_{\beta}\;
\frac{1}{n}\sum_{i=1}^{n} \Big( y_i - \sum_{j=1}^{10} \beta_j z_{ij} \Big)^2 \)
\( \text{Ridge:} \quad \hat{\beta}(\lambda) = \arg\min_{\beta}\;
\text{MSE}(\beta) + \lambda \sum_{j=1}^{10} \beta_j^2 \)
— closed form; smooth shrinkage
\( \text{Lasso:} \quad \hat{\beta}(\lambda) = \arg\min_{\beta}\;
\text{MSE}(\beta) + \lambda \sum_{j=1}^{10} |\beta_j| \)
— coordinate descent (soft-thresholding); exact zeros,
all 10 gone by \(\lambda \approx 10.3\)