Inference Compression: Weight-Matrix Quantization & Pruning (2D)
A 2D weight-matrix heatmap that actually rounds every weight to FP32/FP16/INT8 levels and zeroes the smallest by magnitude, computing real quantization error, memory footprint, matmul latency and speculative-decoding throughput live.
This 2D companion to the layered-network 3D scene works on one concrete weight matrix instead of an abstract node graph. Every one of its 1,232 cells is rounded to real FP32/FP16/INT8 quantization levels computed from the matrix's own value range, and the weakest cells by magnitude are zeroed to hit the chosen sparsity — both operations run on the live data, not a canned animation, so the quantization-error, memory-size and matmul-latency readouts are computed directly from what's on screen. A moving scan column stands in for a matmul sweep, its speed tied to the same latency number, and a token strip below runs real Bernoulli accept/reject trials for speculative decoding at the computed accept rate.
Symmetric linear quantization (scale = max|w| / (2^(bits-1)-1), rounded and dequantized per cell), magnitude-based pruning at a live percentile cutoff, real per-cell MSE, a model-size and matmul-latency estimate derived from the resulting bit-width and density, and a speculative-decoding token strip driven by Bernoulli trials at a computed accept rate.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install