Moderation thresholds

Classifier tuning

Scenario mix

Current post

Tokens scored—
Weighted sum—
Toxicity score—
Decision—

Session stats

Posts scored0
Approved0
Flagged0
Removed0
Approval rate—
Each post is tokenized on whitespace/punctuation, every known token contributes its learned toxicity weight (scaled by sensitivity) to a running sum, and a logistic sigmoid turns that sum into a 0–1 toxicity score. The score is compared against the two thresholds to sort the post into approved / flagged / removed. The top panel highlights exactly which tokens pushed the score up (red) or down (green).