A 3D scatter of image and text embedding vectors that reorganises in real time as contrastive training pulls matching image-caption pairs together and pushes mismatched pairs apart.
Select a batch of image-caption pairs, then press play to step through training iterations and watch the vectors converge, or drag to rotate the embedding space and inspect distances yourself.
Batch selector, training speed slider, play/pause, camera rotate, rebuild embedding space
CLIP's original training set contained about 400 million image-text pairs harvested from the public internet, roughly 400 times larger than the labelled ImageNet dataset it was compared against.
A 3D scatter of image and text embedding vectors that reorganises in real time as contrastive training pulls matching image-caption pairs together and pushes mismatched pairs apart.
A 3D scatter of image and text embedding vectors that reorganises in real time as contrastive training pulls matching image-caption pairs together and pushes mismatched pairs apart.
Select a batch of image-caption pairs, then press play to step through training iterations and watch the vectors converge, or drag to rotate the embedding space and inspect distances yourself.
CLIP's original training set contained about 400 million image-text pairs harvested from the public internet, roughly 400 times larger than the labelled ImageNet dataset it was compared against.