True-goal attractor Decoy attractor Focus agent trail Population (139 agents)

Reward Hacking in Reward-Space: A Goodhart's Law Phase Portrait

The 3D version of this simulator renders a literal reward landscape and lets you watch a single agent climb it. This 2D companion runs the identical hidden dynamics — the same Gaussian true/proxy reward fields, the same finite-difference gradient ascent plus exploration noise — but never draws a landscape at all. Instead every agent (and a full population of 140 at once) is plotted as a single point in reward space: true reward on one axis, proxy reward on the other. Perfect alignment is the diagonal; Goodhart's Law is the swarm drifting above it. A scrolling strip beneath tracks the population's aligned-vs-hacked split over training time, turning a single anecdote into a statistic.