Why Lasso Zeroes Coefficients: the Penalized Loss

The data predicts mileage (\(\text{mpg}\)) based on two predictors: weight (\(\text{wt}\)) and quarter mile time (\(\text{qsec}\)). The gray ellipses are the \(\text{RSS}\) contours; they never move, because the data never changes. The coloured level sets are the penalty. \(\hat\beta(\lambda)\) is where the two meet.

Lasso: \( \text{RSS}(\beta) + \lambda\big(|\beta_1| + |\beta_2|\big) \)

\(\beta_1\) (wt)
\(\beta_2\) (qsec)

Ridge: \( \text{RSS}(\beta) + \lambda\big(\beta_1^2 + \beta_2^2\big) \)

\(\beta_1\) (wt)
\(\beta_2\) (qsec)
\(\text{RSS}\) contours; fixed, they do not depend on \(\lambda\) the \(\text{RSS}\) contour the solution sits on
nested level sets of the penalty; they contract as \(\lambda\) tightens the penalty level set it sits on
 ● \(\hat\beta(\lambda)\)  ★ OLS optimum \(= \hat\beta(0)\) solution path as \(\lambda\) sweeps \(0 \to \lambda_{\max}\)
Penalty strength
\( \definecolor{lassoc}{RGB}{213,94,0}\definecolor{ridgec}{RGB}{48,112,183} \hat{\beta}^{\,\text{lasso}}(\lambda) \;=\; \arg\min_{\beta}\; \text{RSS}(\beta) \;+\; \lambda\,\textcolor{lassoc}{\big(|\beta_1| + |\beta_2|\big)} \)
\( \definecolor{ridgec}{RGB}{48,112,183} \hat{\beta}^{\,\text{ridge}}(\lambda) \;=\; \arg\min_{\beta}\; \text{RSS}(\beta) \;+\; \lambda\,\textcolor{ridgec}{\big(\beta_1^2 + \beta_2^2\big)} \)
\( \text{RSS}(\beta) = \frac{1}{n}\sum_{i=1}^{n} \big(y_i - \beta_1 z_{i1} - \beta_2 z_{i2}\big)^2 \),