Bar height = π(a) Ring = reference π_ref(a)
⚠ Couldn't load the 3D engineThree.js failed to load from the CDN. Check your connection and reload.

RLHF Reward Model: KL-Penalized Alignment

Reinforcement learning from human feedback trains a policy against a learned reward model, but reward alone is dangerous: the policy will happily exploit any gap between what the reward model scores and what a human actually wants. This simulator makes the fix visible — eight candidate responses sit as bars on a 3D stage, each with its own reward value, and a policy distribution climbs those rewards step by step while a KL-divergence penalty of adjustable strength β pulls it back toward a fixed reference policy. Trigger a simulated reward hack to watch a spiked, exploitable reward try to pull all probability mass onto one bad response, and tune β and the learning rate live to see exactly how much anchoring it takes to keep the policy aligned instead of collapsing.