The policy is trained the way real robot-manipulation policies are: by behavior cloning on recorded expert demonstrations, not by hand-coded kinematics. An idealized expert walks the gripper from random start points inside a training region to the fixed target and its (state, action) pairs are stored:
dataset = { (s_i, a_i) } a_i = normalize(target − s_i)
At test time the policy is a locally-weighted k-NN regression over that dataset — the standard non-parametric form of behavior cloning:
w_i = 1 / (‖s − s_i‖ + ε)
π(s) = normalize( Σ w_i a_i / Σ w_i ), over the k nearest s_i
- Demonstrations — more recorded trajectories densely cover the training disk, giving the regression accurate nearby examples anywhere inside it.
- Start distribution shift — pushes the rollout's start point outside the region the demonstrations ever visited. This is covariate shift: the policy must extrapolate from distant, less relevant demonstrations, its output direction gets noisier, the gripper drifts even further from covered states, and the error compounds step after step — the classic behavior-cloning failure mode first analyzed by Ross & Bagnell (DAgger, 2011).
- Execution noise — Gaussian noise on the commanded action, standing in for real actuator/sensor imperfection.
- Nearest demo distance readout — how far the current state sits from anything the policy was actually trained on; it is the live early-warning signal for exactly when a cloned policy is about to fail.
This top-down 2D view shows the same workspace the 3D version drives a two-link SCARA-style arm through: everything happens in a flat plane, so nothing about the underlying policy or physics is lost by dropping the third dimension. Drag to pan the view, scroll or pinch to zoom.