🧪 Batch Size Impact
Interactive SGD landscape: watch how batch size controls gradient-noise magnitude (sigma over sqrt B), steering training into a sharp, low-loss minimum or a flatter, better-generalizing one.
This simulation grounds the batch-size hyperparameter in the actual math of mini-batch stochastic gradient descent: a mini-batch gradient is an average of B per-sample gradients, so its noise variance shrinks as 1/B. A ball descends a real two-well loss landscape under that noise — small batches keep enough noise to escape a narrow sharp minimum and settle into a wide flat one, while large batches converge precisely into whichever minimum is nearest, even a sharp one with lower training loss but a larger train/test generalization gap. Adjust batch size, learning rate and playback speed to watch the trade-off play out step by step.
This simulation investigates the effect of varying batch sizes during machine learning model training. It demonstrates how these choices influence convergence and overall performance.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install