Home▸AI Health Literacy & Translation Tools▸AI-Generated Patient Education Material Accuracy Simulator

🗣 AI-Generated Patient Education Material Accuracy Simulator

This simulation verifies the medical accuracy of educational materials generated by AI.

AI Health Literacy & Translation Tools2DModerate60 FPS
ai-generated-patient-education-accuracy-simulator ↗ Open standalone

AI Drafts the Patient Handout

A language model writes plain-language health guidance from a clinical prompt.

  • ~10: Draft length (claim sentences per handout)
  • 6th grade: Reading level target (plain-language standard)
  • <5 sec: Generation time (per handout draft)
  • Variable: Hallucination risk (depends on model confidence)

Why AI drafts patient materials

Clinics need fast, personalized handouts at scale.

The accuracy problem

Fluent text can still contain false medical claims.

Confidence vs correctness

A confident-sounding sentence is not a verified one.

Low generation confidence raises hallucination risk before any check runs.

Claims Meet the Knowledge Base

Each sentence is compared against a curated, source-linked medical reference set.

  • Curated: Reference sources (guideline-backed medical corpus)
  • ~200 ms: Check latency (per claim comparison)
  • Semantic: Match method (claim-to-source retrieval)
  • Lenient–Strict: Strictness range (adjustable review threshold)

What the knowledge base holds

Vetted clinical facts, dosages, and guideline statements.

How claims are compared

Each sentence is matched against the closest verified source.

Strictness trade-off

Stricter checking narrows what counts as a confirmed match.

Stricter fact-checking filters more borderline claims before scoring.

Every Claim Gets a Number

A 0–100 accuracy score is assigned to each claim after source comparison.

  • 0–100: Score scale (per-claim accuracy rating)
  • ≥70: Pass threshold (typical release cutoff)
  • Source overlap: Scoring basis (plus contradiction detection)
  • Per-claim badge: Output (shown alongside each sentence)

How scores are built

Source agreement and contradiction signals combine into one score.

Reading the distribution

Most claims cluster high; outliers signal real risk.

Score is not a verdict

Scores feed the flagging decision, not the final release.

Low-Confidence Claims Are Flagged

Claims scoring below threshold are marked for clinician review, not silently fixed.

  • Score <70: Flag trigger (or direct contradiction)
  • Human-in-loop: Review queue (clinician sign-off required)
  • Low: False-flag cost (vs. missed medical error)
  • Dosage, dates: Common flag causes (outdated guidance, overreach)

Why flag instead of auto-fix

Silent correction hides errors from clinician oversight.

What gets flagged most

Numeric dosages and time-sensitive guidance flag most often.

Low AI confidence raises the flagged-claim count sharply.

Review, not rejection

A flag routes a claim to a clinician, not the trash.

Verified Content Reaches the Patient

Only claims that pass fact-checking and scoring are compiled into the final handout.

  • 100% checked: Release gate (no unverified claim ships)
  • Held back: Flagged claims (until clinician reviews)
  • Reported: Overall accuracy (as one composite score)
  • Higher: Patient trust (with visible verification layer)

The release gate

Every shipped sentence has passed the accuracy threshold.

What happens to flags

Held claims wait for clinician edits before re-scoring.

Why the layer matters

A review layer turns fluent text into trustworthy guidance.

Verification, not fluency, is what makes AI health text safe to publish.
⚙ Under the hood

This simulation verifies the medical accuracy of educational materials generated by AI.

CanvasBiomedicine

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)