A real 5-layer fully-connected network (24 units/layer, fixed random weights W ~ N(0, scale²/fan_in)) forward-passes a fresh random mini-batch every step. Each layer computes a pre-activation z = Wx + b. With BN off, z goes straight into tanh; with BN on, the batch statistics are used first:
μ_B = mean(z) over the batch
σ²_B = var(z) over the batch
ẑ = (z − μ_B) / √(σ²_B + ε)
y = γ·ẑ + β
a = tanh(y)
Raising the weight-init scale makes each layer's pre-activation variance multiply layer-to-layer (real internal covariate shift) — without BN the histograms spread out or collapse into tanh's saturated tails within a few layers, and the "saturated units" readout (share of activations with |a| > 0.95, where the gradient is nearly zero) climbs. With BN on, γ and β are the network's own learnable affine parameters: γ rescales the normalized distribution's width, β re-centers it, and the per-layer histograms stay a stable width regardless of weight scale or how many layers deep you look.
- Histograms — the distribution of that layer's post-activation values across the whole mini-batch (all units × all samples).
- Variance ratio — layer 5's pre-activation variance divided by layer 1's; near 1× is stable, large values mean the signal is exploding through depth.
- Shift index — standard deviation of the five layers' mean pre-activations; a live measure of how much the "internal covariate shift" the 2015 BN paper named is actually happening right now.