Loading vector space…

Word Embeddings — 3D Vector Space

25 toy word vectors, projected from 8 dimensions to 3D by PCA. Semantically related words cluster together. Click any word, or type one below.

Live Metrics

Vocabulary–
PCA variance kept–
Top match cos θ–
Attention entropy–

Query

Drag to orbit · click a sphere to inspect it

Attention mechanism

Click a token above to make it the query
✕

How it works

Every word here is a hand-built 8-dimensional vector: 4 dims are a one-hot cluster indicator (animal / emotion / number / nature, scaled up so categories separate cleanly) and 4 are shared semantic attributes (intensity, size, positivity, animacy) — this is a toy stand-in for vectors a model like word2vec/GloVe would learn from a text corpus.
cos(a,b) = (a·b) / (‖a‖‖b‖) Clicking a word ranks every other word by this cosine similarity and draws lines to the top-k.
The 3D positions are not the raw 8D vectors — they're the top-3 principal components, found by power-iteration eigen-decomposition of the embedding covariance matrix, so distances you see approximate distances in the full space.
Typing a word not in the vocabulary hashes the text into a pseudo-random 8D vector — a stand-in for an untrained embedding. Its neighbors will look arbitrary, which is the point: without training data a vector carries no meaning.

Attention panel: each token gets a small fixed vector; for a chosen query token i the weight on token j is score(i,j) = (qᵢ·kⱼ)/√d wⱼ = softmax_j( score(i,j) / T ) where T is the temperature slider — lower T sharpens attention onto fewer tokens, higher T spreads it out. Line thickness = weight. Entropy H = −Σ wⱼ·log₂(wⱼ) measures how spread the attention is (0 = all on one token, higher = more diffuse).