HomeAI & Machine LearningRLHF Reward Model: KL-Penalized Alignment

RLHF Reward Model: KL-Penalized Alignment

Interactive 3D RLHF simulator: watch a policy over discrete responses climb a learned reward signal while a KL-divergence penalty anchors it to a reference model, live reward hacking demo included.

AI & Machine Learning3DAdvanced60 FPS📱 Mobile-adapted
navchannia-z-pidkriplenniam-ta-uzhodzhennia-z-liudynoiu-rl-r-explained ↗ Open standalone

Reinforcement learning from human feedback trains a policy against a learned reward model, but reward alone is dangerous: the policy will happily exploit any gap between what the reward model scores and what a human actually wants. This simulator makes the fix visible — eight candidate responses sit as bars on a 3D stage, each with its own reward value, and a policy distribution climbs those rewards step by step while a KL-divergence penalty of adjustable strength β pulls it back toward a fixed reference policy. Trigger a simulated reward hack to watch a spiked, exploitable reward try to pull all probability mass onto one bad response, and tune β and the learning rate live to see exactly how much anchoring it takes to keep the policy aligned instead of collapsing.

⚙ Under the hood

Watch a policy over eight discrete responses climb a learned reward signal in real time while a KL-divergence penalty anchors it to a reference model, with a live reward-hacking demo you can trigger.

reinforcement-learningrlhfalignmentkl-divergencereward-modelpolicy-optimization

3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)