Descent trajectory
Global minimum
Local minima
This 2D companion shows directly, on a flat canvas, why initial values are a hyperparameter in their own right: a hand-defined non-convex loss surface with a shallow local minimum, a medium local minimum and a deeper global minimum sits behind a real momentum gradient-descent optimizer running the update θ ← θ + β·v − lr·∇L(θ) on the surface's analytic gradient every frame. Drag the initial point and the same optimizer settles into a different basin; push the learning rate past the curvature of a well and the same starting point diverges instead of converging; drop it too low and convergence crawls — three real, verifiable effects of the numbers you set before training even starts.