HomeRobotics & KinematicsRobot Language Disambiguation 2D: Ask vs. Act

Robot Language Disambiguation 2D: Ask vs. Act

Interactive 2D top-down simulator of the ask-vs-act policy inside language-guided robots: a vision-language grounding confidence score is computed for every candidate object on the table, and the robot either grasps the top match or asks a clarifying question when the confidence gap is too close to call. Drag to pan the table view, scroll to zoom.

Robotics & Kinematics2DModerate60 FPS📱 Mobile-adapted⇄ 3D version
2d-robotics-topic-24 ↗ Open standalone

Vision-Language-Action robots don't just execute commands — they must first ground a referring phrase like "pick up the red block" in the objects they actually see, and decide whether they're confident enough to act. This top-down 2D simulator renders a table of candidate objects and a robot arm, computes a live grounding-confidence score for every object against the instruction's revealed descriptors, and applies a threshold-gated ask-vs-act policy: when the confidence gap between the top two candidates is wide enough the robot grasps immediately, and when it's too close to call the robot asks a clarifying question that reveals one more descriptor before trying again. Drag the table view to pan and scroll to zoom, and watch the descriptor match matrix light up as each attribute is revealed. Tune the threshold, the number of descriptors given upfront, and the candidate count to see how referential ambiguity trades off against dialogue latency — the exact confidence-gating problem behind real VLA policies like RT-2 and OpenVLA.

⚙ Under the hood

Interactive 2D top-down simulator of the ask-vs-act policy inside language-guided robots: a vision-language grounding confidence score is computed for every candidate object on the table, and the robot either grasps the top match or asks a clarifying question when the confidence gap is too close to call. Drag to pan the table view, scroll to zoom, and watch the descriptor match matrix light up as each attribute is revealed.

roboticsVLAlanguage groundingmanipulationAIhuman-robot interaction

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)