Batch Normalization: Layer-by-Layer Activation Statistics (2D)
2D batch-normalization lab: a real 5-layer network forward-passes a fresh random mini-batch every step and computes each layer's true batch mean, variance and normalized output — toggle BN on/off, push the weight-init scale to force drift, and tune batch size, γ and β to watch the actual formula respond.
The 2015 batch normalization paper's core claim was that deep networks train badly not just because gradients vanish or explode, but because each layer's input distribution keeps shifting as the layers below it update — "internal covariate shift." This 2D companion makes that claim checkable rather than illustrative: a real 5-layer, 24-unit-wide network forward-passes a fresh random mini-batch through fixed weights every training step, computing each layer's genuine batch mean and variance from the actual activations rather than a scripted animation. Push the weight-init scale up and, with normalization off, the five per-layer histograms visibly spread out or collapse into tanh's saturated tails within a few layers — the "saturated units" readout climbs as the gradient there goes flat. Turn batch normalization on and the same weights produce five histograms of comparable width regardless of depth, because each layer's pre-activation is centered and rescaled by its own batch statistics before the learnable γ (scale) and β (shift) affine parameters are reapplied — sliders you control directly, so you can watch normalization's own affine transform undo or preserve the calibration in real time.
A real 5-layer, 24-unit network forward-passes a fresh mini-batch every step; per-layer histograms and a variance-by-depth chart are computed from the actual batch statistics, not a scripted animation. Toggle batch normalization, weight-init scale, batch size, γ and β and watch the formula μ_B, σ²_B, ẑ=(z−μ_B)/√(σ²_B+ε), y=γẑ+β respond live.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install