HomeAI & Machine LearningTransformer Attention Mechanism Explained

🧠 Transformer Attention Mechanism Explained

3D visualization of query-key-value attention heads connecting tokens in a sentence, with sliders for number of heads and softmax temperature and toggles to highlight how attention weights flow between words.

AI & Machine Learning3DModerate60 FPS
transformer-attention-mechanism-explained-lab ↗ Open standalone

Six tokens sit in a ring; glowing beams show how much each token's query attends to every other token's key, with brightness and thickness mapped to the softmax attention weight for each head.

🔬 What It Demonstrates

Scaled dot-product attention: scores from Q·K are divided by √d, passed through a temperature-scaled softmax, and used to blend value vectors from every token into the selected query token.

🎮 How to Use

Pick a query token and how many heads to compute. Lower the temperature to see attention sharpen onto one word, or raise it to flatten it toward uniform. Isolate a single head or overlay them all, and watch value particles flow in along each beam.

💡 Did You Know?

Multi-head attention lets a transformer look for several kinds of relationships at once — one head might track subject-verb agreement while another tracks nearby adjectives, all computed in parallel.

⚙ Under the hood

3D visualization of query-key-value attention heads connecting tokens in a sentence, with sliders for number of heads and softmax temperature and toggles to highlight how attention weights flow between words.

transformersattention-mechanismdeep-learningnlpneural-networksself-attentionmachine-learning

3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)