Supervised boundary Unsupervised (k-means) Reinforcement (TD agent)

Supervised, Unsupervised and Reinforcement Learning Compared (2D)

This 2D companion trains the same three learning paradigms as the 3D version — a hill-climbed decision boundary, Lloyd's k-means clustering and a TD(0) value-learning grid agent — through a plain three-pane canvas view built for reading the mechanics rather than orbiting a scene. A shared control panel drives training speed, label noise, reward sparsity and dataset shape for all three panels at once, while a live readout tracks each paradigm's own metric — classification accuracy, cluster inertia and average reward — epoch by epoch, so the very different information each method needs (full labels, no labels, or only a reward signal) becomes something you can watch converge side by side instead of read about.