Vision-Language-Action (VLA) Robot Simulator
Interactive 3D demonstration of a Vision-Language-Action (VLA) model: a vision encoder identifies scene objects, a language encoder parses a natural-language command into a target and an action, and an action decoder generates the pick-and-place trajectory a robot arm executes end to end.
A minimal 3D stand-in for a Vision-Language-Action model: a vision encoder locates the objects on the table, a language encoder parses a typed command into a target object and an action, and an action decoder plans the pick-and-place trajectory the arm then carries out — all in one continuous pass, without any hand-coded per-object rules.
Interactive 3D demo of a Vision-Language-Action (VLA) model: a vision encoder locates scene objects, a language encoder parses a typed command into a target object and an action, and an action decoder plans the pick-and-place trajectory a robot arm executes end to end.
3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install