The background is a top-down heatmap of the chosen test function, contour-shaded so darker equals lower loss. Each coloured dot descends it using its real update rule: SGD steps opposite the raw gradient (x ← x − lr·∇f); Momentum accumulates a velocity (v ← β·v + lr·∇f, x ← x − v); RMSprop scales each axis by a running average of squared gradients; Adam combines both with bias-corrected first/second moments (β₂ fixed at 0.999).
- Trail — the last 300 steps of each active optimiser's path.
- ★ — the function's known global minimum.
- Gradient noise — random perturbation proportional to |∇f|, mimicking mini-batch stochasticity.