Each of N tissue patches has one true label (tumor / normal). n independent raters each label every patch, with an error probability that can be symmetric (either kind of mistake) or systematically biased (a "liberal" rater only over-calls tumor, a "conservative" rater only under-calls it) — the same asymmetry seen between real pathologists.
Agreement among raters uses Fleiss' kappa, which corrects raw agreement for the agreement expected by chance:
P_i = 1/(n(n-1)) · (Σ_j n_ij² − n)
P̄ = mean_i(P_i) (observed agreement)
p_j = 1/(N·n) · Σ_i n_ij
P̄_e = Σ_j p_j² (chance agreement)
κ = (P̄ − P̄_e) / (1 − P̄_e)
The consensus mask calls a patch "tumor" when the fraction of raters agreeing meets the threshold slider (majority vote at 50%, unanimous at 100%). Its accuracy against the true mask is scored with the Dice coefficient:
Dice = 2·|Consensus ∩ Truth| / (|Consensus| + |Truth|)
- Annotators — how many independent raters label the slide.
- Error rate — each rater's per-patch mistake probability.
- Consensus threshold — how many raters must agree before a patch enters the training ground truth.
- Bias — toggles whether errors are random or systematically directional per rater, which is what drives κ down even when raw agreement looks high.
- Stack tilt / drag — purely a viewing angle on the same layer stack (ground truth, each rater, consensus) drawn top-down when tilt is 0% and as a rotatable isometric stack as tilt increases; it never changes the underlying data.
Real-world relevance: a digital-pathology AI model is only as good as the ground truth it is trained on — low inter-rater κ means noisy labels, which caps how well any downstream model can ever validate against real disease, regardless of the model architecture.