This mirrors the pipeline a real headset or phone (Vision Pro, Quest, HoloLens) runs to turn a camera image into a usable hand skeleton for spatial-computing input, viewed here in 2D — a front-on projection of the hand plane (drag to pan, scroll to zoom):
- 1. Sparse detection. A depth/RGB camera only detects fingertip positions directly, and only with noise — typically a few millimetres of frame-to-frame jitter even under good lighting.
- 2. Temporal smoothing. Each raw detection is passed through an exponential moving average before use, trading a little latency for a lot less jitter:
filtered ← filtered + α·(raw − filtered)
- Low α (heavy smoothing) removes jitter but lags behind fast motion; α → 1 tracks instantly but passes sensor noise straight through — try both sliders together.
- 3. Inverse kinematics. The full skeleton (5 chains, 19 joints, fixed bone lengths) is reconstructed from just the 5 filtered fingertip targets using FABRIK (Forward And Backward Reaching Inverse Kinematics): reach from the tip backward enforcing bone lengths, then from the root forward, iterating a few times per frame until the chain converges.
- 4. Gesture classification. A downstream classifier reads the solved joint geometry — thumb-to-index tip distance for a pinch, relative finger extension for a point or fist — completely independent of which target gesture you selected, exactly as a real gesture recognizer works off the reconstructed skeleton, not the raw camera frame.
When the target is farther from a finger's base than its total bone length, FABRIK can't reach it and instead stretches the chain fully straight toward the target — watch the "IK reach error" readout jump when that happens.