HomeRobotics & KinematicsVision-Language-Action (VLA) Robot Simulator

Vision-Language-Action (VLA) Robot Simulator

Interactive 3D demonstration of a Vision-Language-Action (VLA) model: a vision encoder identifies scene objects, a language encoder parses a natural-language command into a target and an action, and an action decoder generates the pick-and-place trajectory a robot arm executes end to end.

Robotics & Kinematics3DModerate60 FPS
vision-language-action-vla-robot-simulator ↗ Open standalone

A minimal 3D stand-in for a Vision-Language-Action model: a vision encoder locates the objects on the table, a language encoder parses a typed command into a target object and an action, and an action decoder plans the pick-and-place trajectory the arm then carries out — all in one continuous pass, without any hand-coded per-object rules.

⚙ Under the hood

Interactive 3D demo of a Vision-Language-Action (VLA) model: a vision encoder locates scene objects, a language encoder parses a typed command into a target object and an action, and an action decoder plans the pick-and-place trajectory a robot arm executes end to end.

roboticsAIVLAmanipulatorThree.js

3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)