Image-token embedding Text-token embedding Aligned pair link
⚠ Couldn't load the 3D engineThree.js failed to load from the CDN. Check your connection and reload.

3D Multimodal Embedding Space Simulator

Beyond attending across modalities token-by-token, multimodal LLMs are trained so that images and text sharing a concept land close together in one shared vector space — the geometry that makes cross-modal retrieval and grounding possible. This simulation renders that space in 3D: image-token points and text-token points drift and jitter under embedding noise while a spring proportional to alignment strength pulls each true concept pair together, visualizing convergence the way a contrastive training objective would over time. Rotate and zoom to inspect how pairs cluster as alignment strength rises and noise falls.