High-risk candidate Aligned candidate Aligned region (target)
⚠ Couldn't load the 3D engineThree.js failed to load from the CDN. Check your connection and reload.

Constitutional AI: Self-Critique & Revision Loop

A cloud of candidate LLM responses lives in a 3D space of harm, hallucination and unhelpfulness. Each step, an AI critic checks every candidate against whichever written principles are switched on and nudges it toward the aligned region — the same self-critique-and-revise mechanism Constitutional AI uses to build a preference dataset without human harm labels. Toggle principles on and off to see the real trade-off: harmlessness alone over-refuses, helpfulness alone drifts back toward harm and fabrication, and only a balanced constitution converges candidates into the green aligned region.