K-Fold Cross-Validation: Scored on Runs the Fit Never Saw

This dataset is of an experiment on an internal combustion engine (like the one in a car). We are trying to predict a pollutant \(y=\) NOx (the concentration of nitrogen oxides in the exhaust) based on a measure of how rich or lean the fuel–air mixture is, \(x=\) E. K-fold cross-validation splits the runs at random into \(K\) folds; for each fold it fits on the other \(K-1\) and scores the fold it never saw; the \(K\) scores are averaged.

\(\hat f^{(-k)}\) is the degree-\(d\) polynomial fitted without fold \(k\).
\(K\): degree \(d\):

1. One fold at a time

\(x\) (\(E\), centered and divided by its sd)

2. The K held-out errors

fold
CV\(_K\), the mean of the bars  training MSE (fit on all runs)

3. Which degree?

polynomial degree \(d\)
CV\(_K\)  training MSE  one fold's MSE\(_k\)  smallest CV