Home▸LLM Hallucination Detection & Safety▸Clinical AI Uncertainty Communication Simulator

⚠️ Clinical AI Uncertainty Communication Simulator

A method for communicating the level of uncertainty in AI clinical model recommendations to physicians/patients.

LLM Hallucination Detection & Safety2DModerate60 FPS
clinical-ai-uncertainty-communication-simulator ↗ Open standalone

Every AI Answer Carries an Internal Confidence Estimate

Placeholder lead: the model always knows more than it shows.

  • Hidden: Internal estimate (not shown to reviewer by default)
  • 5–98%: Confidence range (varies per query)
  • Fixed tone: Answer text (independent of confidence)
  • Text only: Reviewer sees (no confidence cue yet)

Where the confidence number actually comes from

Placeholder: token probabilities and retrieval agreement set the estimate.

Uniform Confidence Language Regardless of the True Estimate

Placeholder lead: everything sounds equally certain, so nothing stands out.

  • ~88% always: Displayed confidence (flat regardless of truth)
  • Rare: Hedge language (even on weak answers)
  • None: Flag shown (no low-confidence marker)
  • Absent: Reviewer signal (nothing to scrutinize against)

Why flat confidence language is misleading

Placeholder: constant certainty erases the one cue reviewers need.

Displayed Confidence Now Tracks the True Internal Estimate

Placeholder lead: the bar, the number, and the hedge all agree.

  • = true value: Displayed confidence (tight tracking, low gap)
  • Shown <60%: Low-confidence flag (amber/red marker appears)
  • Present: Hedge phrases (scaled to uncertainty)
  • ≈0 pp: Calibration gap (display matches truth)

What a calibrated confidence cue looks like

Placeholder: percentage, color, and wording move together.

Reviewer Attention Differs Sharply Between the Two Styles

Placeholder lead: flags only help if reviewers actually notice them.

  • Flat ~1.0×: Poor style scrutiny (same for all confidence levels)
  • Up to 3.2×: Well style scrutiny (rises as confidence drops)
  • Visible flag: Attention driver (not the true value itself)
  • High (poor): Blind spot risk (low-confidence answers unchecked)

Why scrutiny tracks the display, not the truth

Placeholder: reviewers can only act on what they can see.

Calibrated Signals Cut the Chance a Wrong Answer Slips Through

Placeholder lead: visible honesty about doubt changes review outcomes.

  • Fewer: Unchallenged low-conf errors (with calibrated flags)
  • Higher: Reviewer trust (signal now means something)
  • Reduced: False confidence cost (fewer missed weak answers)
  • Safer review: Net effect (scrutiny matches actual risk)

From honest signals to safer clinical review

Placeholder: small display changes shift where human attention lands.

Placeholder highlight: calibration is a communication problem, not just a modeling one.
⚙ Under the hood

A method for communicating the level of uncertainty in AI clinical model recommendations to physicians/patients.

CanvasBiomedicine

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)