Optimizer trajectory Global minimum Loss / grad-norm charts

Convergence Analysis in Hyperparameter Optimization (2D)

This 2D companion drives the identical non-convex loss surface and the identical SGD / momentum / Adam update rules as the 3D lab, read from directly above instead of orbited: a heatmap shows the basin and its shallow local minima at a glance, the amber trajectory traces every gradient step live, and two log-scale charts — loss vs. iteration and gradient norm vs. iteration — make the stopping criterion ε directly legible as the line the gradient-norm curve has to cross before the run is declared converged.