The heatmap is the exact loss surface from the 3D lab, viewed from directly above instead of orbited: a broad quadratic basin centred on the global minimum plus several Gaussian bumps that create shallow local minima and saddle regions, the same stand-in for a two-parameter slice of a real hyperparameter search space.
L(a,b) = 0.11·[(a−aₘ)² + (b−bₘ)²] + Σᵢ Aᵢ·exp(−[(a−cxᵢ)²+(b−czᵢ)²] / 2sᵢ²)
SGD: a ← a − η·g
Momentum: v ← μ·v − η·g ; a ← a + v
Adam: m,v ← bias-corrected 1st/2nd moment of g ; a ← a − η·m̂/(√v̂+ε)
Each frame the marker takes one gradient step on that surface. Learning rate sets step size, momentum carries velocity through narrow valleys, gradient noise mimics mini-batch stochasticity, and the run is declared converged once the gradient norm ‖∇‖ stays below the threshold ε for several consecutive steps — exactly the criterion real training loops use to stop early.
- Left pane — the loss heatmap (dark = low loss) with the live trajectory (amber trail) and the global minimum marked by a green ring.
- Top-right chart — loss vs. iteration, log scale, the standard way convergence rate is read in the optimization literature.
- Bottom-right chart — gradient norm vs. iteration, log scale, with the ε stopping line drawn where it is actually compared.