This is a real, running t-SNE optimisation — not a canned animation. A synthetic dataset lives in 5 high-dimensional coordinates; every frame, gradient descent moves each point's 2D position so that its neighbour probabilities in the map match its neighbour probabilities in the original 5D space:
p(j|i) ∝ exp(-||xi-xj||² · beta_i) (high-D, Gaussian, per-point beta solved for target perplexity)
q(j|i) ∝ (1 + ||yi-yj||²)⁻¹ (2D map, Student-t, heavy tails)
∂KL/∂yi = 4 Σj (p_ij - q_ij) · qNum_ij · (yi - yj)
The neighbor-affinity graph draws a line between two points only when their symmetrized high-D affinity Pij clears a threshold — so you are looking directly at which pairs the optimiser is trying to pull together, in the plane, without a 3D camera in the way.
- Perplexity — target effective neighbour count per point; low = tight local cliques, high = broad, blurred structure that can fuse separate clusters.
- Cluster separation — how far apart the ground-truth clusters sit in 5D before t-SNE ever sees them.
- Clusters — number of ground-truth groups in the synthetic dataset (3–7).
- Learning rate — the gradient-descent step size (η); too low crawls, too high can oscillate before settling.
Because t-SNE only preserves local neighbour relationships, cluster sizes and the gaps between them on the final map carry no reliable meaning — a caveat popularised by "How to Use t-SNE Effectively" (Wattenberg, Viégas & Johnson, 2016).