← All STAT 413 applets

Loss Surfaces: OLS vs Ridge vs Lasso

All three estimators minimize a criterion over the two coefficients \((\beta_1, \beta_2)\), computed here from a real dataset (you can switch between datasets if you want). Toggle between ridge and lasso, and drag any surface to rotate all three together; The contour plot below each surface is the same function seen from above.

Data
\(x_1\) = , \(x_2\) = , \(y\) = ; \(\operatorname{corr}(x_1,x_2)\) =
Penalty
at \(\lambda = 0\) the penalty vanishes and the sum equals OLS
Zoom

\( \definecolor{ridgec}{RGB}{25,118,210}\definecolor{lassoc}{RGB}{230,81,0} \mathrm{RSS}(\beta_1,\beta_2) = \sum_{i=1}^{n} \big(y_i - \beta_1 x_{i1} - \beta_2 x_{i2}\big)^2 \)  (predictors and response standardized).

Color encodes the criterion's height; all three panels share one fixed vertical scale (the sum's range at the maximum \(\lambda\)), so the heights literally add — OLS + penalty = sum — the penalty visibly grows with \(\lambda\), and the OLS panel never changes. ● marks each criterion's minimizer (on both the surface and the contour); ✕ marks the OLS minimizer for reference.

OLS

\( \mathrm{RSS}(\beta_1,\beta_2) \)
\(\beta_1\) \(\beta_2\)
\(\hat\beta\) =
+

Penalty

\( \textcolor{ridgec}{\lambda\,(\beta_1^2+\beta_2^2)} \)\( \textcolor{lassoc}{\lambda\,(|\beta_1|+|\beta_2|)} \)
\(\beta_1\) \(\beta_2\)
minimized at
=

RidgeLasso

\( \mathrm{RSS}(\beta_1,\beta_2) + \textcolor{ridgec}{\lambda\,(\beta_1^2+\beta_2^2)} \)\( \mathrm{RSS}(\beta_1,\beta_2) + \textcolor{lassoc}{\lambda\,(|\beta_1|+|\beta_2|)} \)
\(\beta_1\) \(\beta_2\)
\(\hat\beta\) =