Home▸AI & Machine Learning▸RLHF Reward Model: KL-Penalized Alignment (2D)

RLHF Reward Model: KL-Penalized Alignment (2D)

2D bar-chart RLHF simulator: watch a policy over eight discrete responses climb a learned reward signal while an adjustable KL-divergence penalty anchors it to a reference distribution, with a live reward-hacking demo.

AI & Machine Learning2DAdvanced60 FPS📱 Mobile-adapted⇄ 3D version
2d-navchannia-z-pidkriplenniam-ta-uzhodzhennia-z-liudynoiu-rl-r-explained ↗ Open standalone

Reinforcement learning from human feedback trains a policy against a learned reward model, but reward alone is dangerous: the policy will happily exploit any gap between what the reward model scores and what a human actually wants. This 2D companion makes the fix visible as a live bar chart — eight candidate responses each carry their own reward value, and a policy distribution climbs those rewards step by step while a KL-divergence penalty of adjustable strength β pulls it back toward a fixed reference policy drawn as a dashed tick on every bar. Trigger a simulated reward hack to watch a spiked, exploitable reward try to pull all probability mass onto one bad response, and tune β and the learning rate live to see exactly how much anchoring it takes to keep the policy aligned instead of collapsing.

⚙ Under the hood

2D bar-chart RLHF simulator: watch a policy over eight discrete responses climb a learned reward signal while an adjustable KL-divergence penalty anchors it to a reference distribution, with a live reward-hacking demo.

reinforcement-learningrlhfalignmentkl-divergencereward-modelpolicy-optimization

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)