This is a weight-level companion to the layered-network 3D view: instead of watching abstract signal pulses travel through a node graph, you watch one real weight matrix (28×44 values, a stand-in slice of a linear layer) get quantized and pruned cell by cell, with the resulting error, size and speed read straight off that matrix rather than from a fixed formula alone.
scale = max|w| / (2^(bits-1) - 1) // symmetric linear quantization
w_q = round(w / scale) · scale
MSE = mean((w - w_q)^2) over unpruned cells
density = 1 - sparsity
size = params · bits/8 · density
latency = L0 · k(precision) · (1 - 0.55·sparsity)
E[tok/rd] = (1 − α^(γ+1)) / (1 − α)
- Quantization — every cell is rounded to the nearest of 2^bits evenly-spaced levels spanning the matrix's own value range; INT8 has only 256 levels so rounding noise (the heatmap's speckle) is visible, FP32 is effectively exact.
- Pruning — cells below the sparsity-th percentile of |w| are zeroed (magnitude pruning, computed on the live matrix each time you move the slider) and rendered as dark cells; above ~50% sparsity it also starts eroding the draft model's accept rate α.
- Systolic scan — the moving highlight column mimics a matmul sweeping across the matrix; its speed is driven by the same computed latency as the stat box, so a slower precision/sparsity setting visibly slows the sweep.
- Draft+verify — a cheap draft model proposes γ tokens; the target model verifies them in one pass, accepting a prefix and rejecting at the first miss, drawn with real Bernoulli trials at the computed accept rate α.
Real serving stacks (TensorRT-LLM and similar) combine exactly this: INT8/FP8 quantization, magnitude or structured pruning, and speculative/lookahead decoding, to cut cost per generated token.