Every stage is computed for real, in order, from the pixels you set:
Conv1: y[o] = b1[o] + Σ conv(x, W1[o]) (K1 kernels, 3×3, same-padding)
ReLU1: max(0, y)
Pool1: 2×2 max-pool → 12×12 → 6×6
Conv2: y[o] = b2[o] + Σᵢ conv(pool1[i], W2[o][i]) (K2 kernels, multi-channel)
ReLU2 → Pool2 (6×6 → 3×3)
Flatten: 3×3×K2 → vector
Dense: logits[c] = b[c] + Σ vec[f]·Wd[c][f]
Softmax: p[c] = eˡᵒᵍᶦᵗˢ⁽ᶜ⁾ / Σ eˡᵒᵍᶦᵗˢ
Conv2 is a genuine multi-channel convolution: each of its output maps is the sum of a separate 3×3 kernel convolved with every Pool1 channel, plus a bias — exactly how a real CNN mixes channels between layers. Nothing here is pre-rendered; changing a kernel-count slider or the "Reseed weights" button regenerates the weight tensors and every downstream feature map, dense logit and softmax probability recomputes instantly.
- Conv1 / Conv2 — raw weighted sums, can be negative.
- ReLU1 / ReLU2 — f(x)=max(0,x) clips negatives to zero, the network's non-linearity.
- Pool1 / Pool2 — 2×2 max-pooling halves each spatial dimension.
- Flatten — all pooled channels concatenated into one vector.
- Softmax — the dense layer's logits turned into class probabilities that sum to 1.
The weights are randomly initialized (Reseed draws a fresh random set), exactly as a real CNN starts before training — so the class it "picks" is architecturally real but not a learned, meaningful recognition. This lab teaches the mechanics of a full forward pass; training would require labeled data and backpropagation, which is outside the scope of an in-browser demo.