Gradient Descent in 1D
\(L(\beta)\) represents a loss function for parameter \(\beta\). The ball represents our current estimate, \(\beta^{(t)}\) and it repeatedly takes the step \(\beta^{(t+1)} = \beta^{(t)} - \rho\, \frac{dL(\beta^{(t)})}{d\beta}\). Drag anywhere on the plot to move
the starting point \(\beta^{(0)}\). Try: crank \(\rho\) past 1 on the smooth bowl; watch the
V-shape bounce forever; drag \(\beta^{(0)}\) across the hump of the double well.
The function, the ball, and its path
\(L(\beta)\)
tangent at \(\beta^{(t)}\) (slope \(\frac{dL(\beta^{(t)})}{d\beta}\))
path so far; dashed = the next step
★ local minima
Is it getting anywhere? — \(L(\beta^{(t)})\) against iteration (dashed line = lowest possible value)
iteration \(t\)
\(L(\beta^{(t)})\)
\(t\) =
| \(\beta^{(t)}\) =
| \(L(\beta^{(t)})\) =
| \(\frac{dL(\beta^{(t)})}{d\beta}\) =
| step \(-\rho\, \frac{dL(\beta^{(t)})}{d\beta}\) =
current function:
update rule: \( \beta^{(t+1)} \;=\; \beta^{(t)} \;-\; \rho\, \frac{dL(\beta^{(t)})}{d\beta} \)