Reliability diagram
Stated confidence (x) vs. actual accuracy (y) per 10%-wide bin. Dot size = number of trials in that bin.
Perfect calibration (diagonal) Your bins Overconfident region Underconfident region

Confidence Calibration Lab

Metacognitive monitoring is the ability to accurately judge how likely your own answer is to be correct. This simulator makes that judgment measurable: before each trial you set a confidence rating, then answer, and the reliability diagram bins every stated confidence against the fraction of trials in that bin that actually turned out correct. A perfectly calibrated forecaster's bins sit on the diagonal; systematic overconfidence pulls them below it, systematic underconfidence pushes them above it. The batch simulator can generate dozens of trials under a chosen bias so the three calibration archetypes — overconfident, underconfident, well-calibrated — are visible immediately, and the running Brier score gives a single number for how good your (or the simulated) confidence judgments are.