🗜️ Inference Optimization & Compression: Quantization, Pruning & Speculative Decoding
Watch a transformer shrink and speed up: switch FP32/FP16/INT8 quantization, prune weak connections, and enable speculative decoding to see model size, latency and throughput change live.
Computer Science3DModerate60 FPS
⚙ Under the hood
Switch a transformer between FP32/FP16/INT8 quantization, prune weak connections and toggle speculative decoding to watch model size, latency and throughput change live.
Three.jsquantizationpruningspeculative decodingLLM inferenceInstancedMesh
3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install