Each dot is a simulated patient positioned by two normalized
features (e.g. biomarker level vs. symptom score). K-means starts
with k random centroids, assigns every point to its nearest
centroid, then moves each centroid toward the mean of its
assigned points — repeated until the centroids stop moving.
assign: cluster(i) = argmin_c ||x_i − centroid_c||²
update: centroid_c ← centroid_c + 0.6·(mean(members) − centroid_c)
repeat until centroid drift ≈ 0
- Sample size — how many simulated patient records are plotted.
- Feature separation — how strongly the underlying risk factor pulls the true groups apart on the plane.
- Cluster count (k) — how many risk groups the algorithm tries to find; picking k badly under- or over-splits real structure.
This is a genuine unsupervised learning technique used in real
healthcare analytics to surface patient risk segments — cohorts a
clinician can then investigate — directly from raw records, without
needing labelled outcomes in advance.