harm × hallucination — drag/scroll
High-risk candidate Aligned candidate Aligned region (target)

Constitutional AI: Self-Critique & Revision Loop (2D)

A scatter of candidate LLM responses lives in a space of harm, hallucination and unhelpfulness. Each step, an AI critic checks every candidate against whichever written principles are switched on and nudges it toward the aligned region — the same self-critique-and-revise mechanism Constitutional AI uses to build a preference dataset without human harm labels. This 2D edition plots harm against hallucination directly, sizes each dot by its unhelpfulness, and adds a trend chart and a live gauge alongside the pannable, zoomable scatter. Toggle principles on and off to see the real trade-off: harmlessness alone over-refuses, helpfulness alone drifts back toward harm and fabrication, and only a balanced constitution converges candidates into the green aligned region.