🧠 Transformer Self-Attention Explorer
Watch self-attention weights emerge between tokens across stacked encoder layers and heads: softmax(QK^T/√d_k) lights up token-to-token connections in real time as temperature reshapes how sharp or diffuse the attention becomes.
AI & Machine Learning3DAdvanced60 FPS
⚙ Under the hood
Watch self-attention weights emerge between tokens across stacked encoder layers and heads: softmax(QK^T/√d_k) lights up token-to-token connections in real time as temperature reshapes how sharp or diffuse the attention becomes.
Three.jsAITransformerAttentionNLPDeep Learning
3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install