Constitutional AI: Self-Critique & Revision Loop
Interactive 3D simulator of Constitutional AI: watch a cloud of candidate LLM responses get critiqued against written principles and iteratively revised toward a harmless, honest, helpful region — with real principle trade-offs like the helpfulness/harmlessness tension.
A cloud of candidate LLM responses lives in a 3D space of harm, hallucination and unhelpfulness. Each step, an AI critic checks every candidate against whichever written principles are switched on and nudges it toward the aligned region — the same self-critique-and-revise mechanism Constitutional AI uses to build a preference dataset without human harm labels. Toggle principles on and off to see the real trade-off: harmlessness alone over-refuses, helpfulness alone drifts back toward harm and fabrication, and only a balanced constitution converges candidates into the green aligned region.
Watch a cloud of candidate LLM responses get critiqued against written constitutional principles and iteratively revised toward a harmless, honest, helpful region — the self-critique-and-revise mechanism behind Constitutional AI and RLAIF.
3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install