system idle
Robot arm (top view)
Red ball
Green cube
Yellow cylinder
Blue box
Orange box
A top-down 2D stand-in for a Vision-Language-Action model: a vision encoder locates the objects on the table, a language encoder parses a typed command into a target object and an action, and an action decoder solves a 2-link inverse-kinematics trajectory the arm then carries out — all in one continuous pass, without any hand-coded per-object rules.