Demonstration coverage Target Gripper trail
Drag to pan · Scroll to zoom

Behavior Cloning & Distribution Shift in Robot Manipulation (2D)

Instead of hand-coding a control law, this simulator trains a manipulation policy the way real learning-based robots are trained: by cloning an expert's demonstrations. Recorded (state, action) pairs from an idealized expert form a training dataset, and a k-nearest-neighbor regression policy reproduces the expert's behavior anywhere near that data. Push the rollout's start point outside the demonstrated region and watch covariate shift take hold live — the policy's commanded direction gets noisier the farther it is from any demonstration, the gripper drifts into even less-covered territory, and the error compounds exactly as it does in real imitation-learning failures, the core problem that motivated on-policy corrections like DAgger and RL fine-tuning. This top-down view renders the identical policy math as a flat 2D scene you can pan and zoom.