HomeAI & Machine LearningVision-Language Models: Image-Text Attention & Grounding

🖼 Vision-Language Models: Image-Text Attention & Grounding

Watch a vision-language model attend across image patches to answer visual questions or generate captions — tune model scale and noise to see grounding accuracy and attention sharpness change live.

AI & Machine Learning3DModerate60 FPS
vision-language-models-image-text-attention-grounding ↗ Open standalone
⚙ Under the hood

Watch a vision-language model attend across image patches to answer visual questions or generate captions — tune model scale and noise to see grounding accuracy and attention sharpness change live, the mechanism behind GPT-4V, Gemini Vision, LLaVA, CLIP and Florence-2.

Three.jsvision-language modelVLMattentionCLIPVQAimage captioningInstancedMesh

3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)