Behavior Cloning & Distribution Shift in Robot Manipulation (2D)
Top-down 2D behavior-cloning manipulation policy: a k-nearest-neighbor regression policy learns a two-link arm's reach from recorded expert demonstrations. Drag to pan, scroll to zoom, and push test rollouts outside the training region to watch covariate shift and compounding error take hold live.
Instead of hand-coding a control law, this simulator trains a manipulation policy the way real learning-based robots are trained: by cloning an expert's demonstrations. Recorded (state, action) pairs from an idealized expert form a training dataset, and a k-nearest-neighbor regression policy reproduces the expert's behavior anywhere near that data. Push the rollout's start point outside the demonstrated region and watch covariate shift take hold live — the policy's commanded direction gets noisier the farther it is from any demonstration, the gripper drifts into even less-covered territory, and the error compounds exactly as it does in real imitation-learning failures, the core problem that motivated on-policy corrections like DAgger and RL fine-tuning. This top-down view renders the identical policy math as a flat 2D scene you can pan and zoom.
Top-down 2D behavior-cloning manipulation policy: a k-nearest-neighbor regression policy learns a two-link arm's reach from recorded expert demonstrations. Drag to pan, scroll to zoom, and push test rollouts outside the training region to watch covariate shift and compounding error take hold live.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install