Underfitting, Good Fit, Overfitting

This dataset is of an experiment using the internal combusion engine (like what cars have). We are trying to predict a pollutant \(y=\) NOx (the concentration of nitrogen oxides in the exhaust) based on a measure of how rich or lean the fuel–air mixture is, \(x=\) E. Slide the degree and watch the two losses: training loss can only fall, validation loss falls and then rises, and the gap between them is what overfitting looks like. The loss axis is fixed at 0–1, so once a loss is off the chart it stays off the chart (the readouts still give it). By degree 43 the polynomial passes through all 44 training runs exactly — training loss 0 — and shoots off between them; the validation loss is then around \(10^{22}\).

degree \(d\):

The fit

\(x\) (\(E\), centered and divided by its sd)
training points (the fit sees these)  validation points (never seen by the fit)  degree-\(d\) fit, colored by regime

Loss vs degree

polynomial degree \(d\)
training loss  validation loss   shaded: the gap   smallest validation loss   off the chart (above 1)
  1. Which is the best value for \(d\) and why?
  2. What is happening when at \(d=44\) to the fit and the training set?
  3. Why does this induce a lot of valdiation error?
  4. How does this relate to bias and variance?