HomeComputer ScienceInference Optimization & Compression: Quantization, Pruning & Speculative Decoding

🗜️ Inference Optimization & Compression: Quantization, Pruning & Speculative Decoding

Watch a transformer shrink and speed up: switch FP32/FP16/INT8 quantization, prune weak connections, and enable speculative decoding to see model size, latency and throughput change live.

Computer Science3DModerate60 FPS
inference-optimization-compression-quantization-pruning ↗ Open standalone
⚙ Under the hood

Switch a transformer between FP32/FP16/INT8 quantization, prune weak connections and toggle speculative decoding to watch model size, latency and throughput change live.

Three.jsquantizationpruningspeculative decodingLLM inferenceInstancedMesh

3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)