Inference Compression: Weight-Matrix Quantization & Pruning (2D)

This 2D companion to the layered-network 3D scene works on one concrete weight matrix instead of an abstract node graph. Every one of its 1,232 cells is rounded to real FP32/FP16/INT8 quantization levels computed from the matrix's own value range, and the weakest cells by magnitude are zeroed to hit the chosen sparsity — both operations run on the live data, not a canned animation, so the quantization-error, memory-size and matmul-latency readouts are computed directly from what's on screen. A moving scan column stands in for a matmul sweep, its speed tied to the same latency number, and a token strip below runs real Bernoulli accept/reject trials for speculative decoding at the computed accept rate.