A synthetic samples × genes expression matrix X is built from K cluster centroids placed randomly in a 10-dimensional gene-expression space, scaled by the separation slider; each sample is a centroid plus Gaussian noise at the chosen σ. The genuine PCA pipeline runs entirely client-side, from scratch:
X̄ = mean(X, axis=0)
Xc = X − X̄ (mean-centred)
C = (1 / (N−1)) · Xcᵀ · Xc (10×10 covariance matrix)
C = V · Λ · Vᵀ (Jacobi eigenvalue algorithm)
PC1, PC2 = the 2 eigenvectors of C with the largest eigenvalues
scores = Xc · [PC1 | PC2] (projection plotted below)
The eigendecomposition uses the classical cyclic Jacobi rotation method: at every step it finds the largest off-diagonal entry of the (symmetric) covariance matrix and applies a rotation that zeroes it, repeating until all off-diagonal entries vanish. The diagonal that remains holds the eigenvalues, and the accumulated rotations give the eigenvectors — no library, no shortcuts. As a correctness check, the sum of all 10 eigenvalues always equals the trace of the covariance matrix (the sum of per-gene variances), shown live above.
- Scatter plot — every sample projected onto PC1 (x) and PC2 (y), coloured by its true generating cluster (a label PCA never sees).
- Scree pane — the top-5 eigenvalues as a bar chart, i.e. how much variance each principal component explains.
- Explained variance — (λ₁+λ₂) / Σλᵢ: raise separation or lower noise and this climbs toward 100%, because the cluster structure becomes more aligned with just 2 directions.