SGD iterate Sharp minimum Flat minimum
⚠ Couldn't load the 3D engineThree.js failed to load from the CDN. Check your connection and reload.

Batch Size Impact

This simulation grounds the batch-size hyperparameter in the actual math of mini-batch stochastic gradient descent: a mini-batch gradient is an average of B per-sample gradients, so its noise variance shrinks as 1/B. A ball descends a real two-well loss landscape under that noise — small batches keep enough noise to escape a narrow sharp minimum and settle into a wide flat one, while large batches converge precisely into whichever minimum is nearest, even a sharp one with lower training loss but a larger train/test generalization gap. Adjust batch size, learning rate and playback speed to watch the trade-off play out step by step.