← All STAT 413 applets

Be the Optimizer

Below is a real dataset: 32 cars road-tested by Motor Trend in 1974. Your job: predict \(\text{mpg}\) from the 10 other variables by choosing the coefficients \(\beta_1, \dots, \beta_{10}\) by hand to make the loss as small as possible. Each predictor is standardized so that no intercept is needed. There is no picture to guide you — only the number. This is exactly the search gradient descent automates.

The data before standardization

Loss to minimize
current loss
optimal loss:  |  your best:  |  slider adjustments: 0
within 2% of the optimum — as good as converged!

The 10 knobs

Careful: the knobs are coupled. The best value for one \(\beta_j\) depends on all the others, so fixing one knob un-fixes the rest.
The solution is only shown — you still have to dial it in yourself.

Your loss history (dashed line = optimum)

\( \text{OLS:} \) \( L_\text{OLS}(\beta_1, \dots, \beta_{10}) \) \( = \frac{1}{n}\sum_{i=1}^{n} \Big( y_i - \sum_{j=1}^{10} \beta_j z_{ij} \Big)^2 \)
\( \text{Ridge:} \) \( L_\text{R}(\beta_1, \dots, \beta_{10}) \) \( = L_\text{OLS}(\beta_1, \dots, \beta_{10}) + \lambda \sum_{j=1}^{10} \beta_j^2, \qquad \lambda = 0.5 \)
\( \text{Lasso:} \) \( L_\text{LASSO}(\beta_1, \dots, \beta_{10}) \) \( = L_\text{OLS}(\beta_1, \dots, \beta_{10}) + \lambda \sum_{j=1}^{10} |\beta_j|, \qquad \lambda = 1 \)