⚠ Couldn't renderCanvas 2D context failed to initialize. Try reloading.

RLHF Reward Model: KL-Penalized Alignment (2D)

Reinforcement learning from human feedback trains a policy against a learned reward model, but reward alone is dangerous: the policy will happily exploit any gap between what the reward model scores and what a human actually wants. This 2D companion makes the fix visible as a live bar chart — eight candidate responses each carry their own reward value, and a policy distribution climbs those rewards step by step while a KL-divergence penalty of adjustable strength β pulls it back toward a fixed reference policy drawn as a dashed tick on every bar. Trigger a simulated reward hack to watch a spiked, exploitable reward try to pull all probability mass onto one bad response, and tune β and the learning rate live to see exactly how much anchoring it takes to keep the policy aligned instead of collapsing.