🖼 Vision-Language Models: Image-Text Attention & Grounding
Watch a vision-language model attend across image patches to answer visual questions or generate captions — tune model scale and noise to see grounding accuracy and attention sharpness change live.
AI & Machine Learning3DModerate60 FPS
⚙ Under the hood
Watch a vision-language model attend across image patches to answer visual questions or generate captions — tune model scale and noise to see grounding accuracy and attention sharpness change live, the mechanism behind GPT-4V, Gemini Vision, LLaVA, CLIP and Florence-2.
Three.jsvision-language modelVLMattentionCLIPVQAimage captioningInstancedMesh
3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install