🗣 AI-Generated Patient Education Material Accuracy Simulator
This simulation verifies the medical accuracy of educational materials generated by AI.
AI Drafts the Patient Handout
A language model writes plain-language health guidance from a clinical prompt.
- ~10: Draft length (claim sentences per handout)
- 6th grade: Reading level target (plain-language standard)
- <5 sec: Generation time (per handout draft)
- Variable: Hallucination risk (depends on model confidence)
Why AI drafts patient materials
Clinics need fast, personalized handouts at scale.
The accuracy problem
Fluent text can still contain false medical claims.
Confidence vs correctness
A confident-sounding sentence is not a verified one.
Low generation confidence raises hallucination risk before any check runs.
Claims Meet the Knowledge Base
Each sentence is compared against a curated, source-linked medical reference set.
- Curated: Reference sources (guideline-backed medical corpus)
- ~200 ms: Check latency (per claim comparison)
- Semantic: Match method (claim-to-source retrieval)
- Lenient–Strict: Strictness range (adjustable review threshold)
What the knowledge base holds
Vetted clinical facts, dosages, and guideline statements.
How claims are compared
Each sentence is matched against the closest verified source.
Strictness trade-off
Stricter checking narrows what counts as a confirmed match.
Stricter fact-checking filters more borderline claims before scoring.
Every Claim Gets a Number
A 0–100 accuracy score is assigned to each claim after source comparison.
- 0–100: Score scale (per-claim accuracy rating)
- ≥70: Pass threshold (typical release cutoff)
- Source overlap: Scoring basis (plus contradiction detection)
- Per-claim badge: Output (shown alongside each sentence)
How scores are built
Source agreement and contradiction signals combine into one score.
Reading the distribution
Most claims cluster high; outliers signal real risk.
Score is not a verdict
Scores feed the flagging decision, not the final release.
Low-Confidence Claims Are Flagged
Claims scoring below threshold are marked for clinician review, not silently fixed.
- Score <70: Flag trigger (or direct contradiction)
- Human-in-loop: Review queue (clinician sign-off required)
- Low: False-flag cost (vs. missed medical error)
- Dosage, dates: Common flag causes (outdated guidance, overreach)
Why flag instead of auto-fix
Silent correction hides errors from clinician oversight.
What gets flagged most
Numeric dosages and time-sensitive guidance flag most often.
Low AI confidence raises the flagged-claim count sharply.
Review, not rejection
A flag routes a claim to a clinician, not the trash.
Verified Content Reaches the Patient
Only claims that pass fact-checking and scoring are compiled into the final handout.
- 100% checked: Release gate (no unverified claim ships)
- Held back: Flagged claims (until clinician reviews)
- Reported: Overall accuracy (as one composite score)
- Higher: Patient trust (with visible verification layer)
The release gate
Every shipped sentence has passed the accuracy threshold.
What happens to flags
Held claims wait for clinician edits before re-scoring.
Why the layer matters
A review layer turns fluent text into trustworthy guidance.
Verification, not fluency, is what makes AI health text safe to publish.
This simulation verifies the medical accuracy of educational materials generated by AI.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install