🩺 Diagnostic Error Root-Cause Analysis Simulator
This simulation focuses on root cause analysis of diagnostic errors. It helps users identify the underlying factors that contribute to misdiagnosis and develop strategies to prevent such errors in clinical practice.
Sentinel Event Identification & the Case for Formal Diagnostic RCA
Diagnostic error is estimated to affect roughly 12 million adult outpatients per year in the United States alone — about 1 in 20 visits — with perhaps half of those errors carrying potential for harm (illustrative estimates from the 2015 National Academies of Medicine report, "Improving Diagnosis in Health Care"). Unlike a wrong-site surgery or medication overdose, a missed or delayed diagnosis rarely announces itself at the moment it occurs. It surfaces later — a returning patient, an autopsy discrepancy, a malpractice claim — which is precisely why a structured trigger and a disciplined root-cause analysis (RCA) process are needed to catch it at all.
- ~12 M: US outpatient diagnostic errors/yr (illustrative, NAM 2015 report)
- ~10–20%: Autopsy-confirmed major discordance (illustrative, pooled autopsy studies)
- ~30%: Diagnostic error share of claims (illustrative, malpractice literature)
- 45 days: Time to convene formal RCA (typical sentinel-event policy window)
What counts as a diagnostic error, and who catches it
The National Academies of Medicine (formerly IOM) defines diagnostic error as "the failure to (a) establish an accurate and timely explanation of the patient's health problem(s) or (b) communicate that explanation to the patient." This definition deliberately includes correct-but-late diagnoses and correctly-made-but-never-communicated diagnoses, not just outright wrong answers.
Most diagnostic errors are discovered through one of a handful of recurring triggers: a "bounce-back" patient returning to the emergency department within 72 hours with a worsened version of the same problem; a discrepancy between the antemortem clinical diagnosis and the autopsy or pathology report; a critical imaging or lab result that was never acted upon (a "closed-loop" failure); a formal malpractice claim; or a voluntary safety-event report filed by a clinician, nurse, or the patient themselves.
Because so few diagnostic errors are self-evident at the time they happen, health systems increasingly use "trigger tools" — algorithms that scan the record for patterns statistically associated with missed diagnoses (e.g., an ED revisit with escalation of care, an unplanned ICU transfer within 24 hours of admission, a test ordered and never resulted) — to proactively surface candidate cases for review rather than waiting for harm to be reported.
A useful mental model: diagnostic error is rarely a single wrong guess. It is far more often a process failure — a breakdown somewhere in the chain of information gathering, interpretation, and communication that stretches from the patient's first symptom to the final treatment decision. RCA exists to find where that chain broke.
Sentinel event policy and the blame-free mandate
Once a case is flagged, most hospitals classify it against a sentinel-event or serious-safety-event policy modeled on Joint Commission standards: an unexpected occurrence involving death or serious physical/psychological injury, or the risk thereof, unrelated to the natural course of the patient's illness. A confirmed or highly suspected diagnostic error resulting in significant harm typically meets this bar and triggers a comprehensive systematic analysis within a defined window — commonly 45 days from event identification to submission of an action plan.
The single most important precondition for a useful RCA is psychological safety: the review must be explicitly blame-free and, wherever legally possible, protected from discovery as peer-review privileged material. If clinicians believe an honest account of "I anchored on the initial read and didn't revisit it" will be used against them individually, the true causal chain never surfaces — the team instead reconstructs a defensible narrative rather than an accurate one. High-reliability organizations treat the individual clinician as the last, most visible link in a much longer chain of latent system conditions (Reason's "Swiss cheese" model), and the RCA charter is written accordingly.
Timeline Reconstruction — Rebuilding the Diagnostic Journey Minute by Minute
Before any causal analysis can begin, the RCA team must establish an objective, minute-by-minute chronology of everything that happened from the patient's first contact with the system to the moment the error was discovered. This timeline is built entirely from the factual record — timestamps, order entries, nursing notes, imaging metadata, phone logs — deliberately stripped of interpretation or blame at this stage.
- 6–10 h: Median ED length of stay reviewed (illustrative, typical RCA case window)
- 8–15: Distinct data sources per case (EHR, PACS, phone logs, staffing records)
- <60 min: Critical-result reporting standard (typical hospital policy for STAT/critical values)
- ~70%: Cases with a documented "trigger" gap (illustrative, RCA case-series estimate)
Assembling the factual timeline
A rigorous diagnostic-error timeline typically pulls from every system the patient touched: the EHR audit trail (who opened which note, and when), order-entry and results systems (when a test was ordered, resulted, and viewed — and by whom), the radiology PACS (image acquisition time vs. preliminary vs. final read time), nurse call logs and vital-sign flowsheets, telephone and paging records for any verbal communication, staffing rosters (census, nurse-to-patient ratio, whether the ordering clinician was mid-shift-change), and, where available, patient-reported recollection of symptoms and instructions given at discharge.
Each event is time-stamped to the minute where possible and entered onto a single shared timeline visible to the whole RCA team. Gaps are marked explicitly rather than filled in with assumption — for example, "critical finding flagged in PACS at 14:32; first evidence of clinician acknowledgment at 19:50" is left as a 5-hour, 18-minute gap rather than assumed to represent any particular cause. Only after the timeline is complete and agreed upon does the team move to interpreting why those gaps occurred.
Distinguishing timeline facts from causal narrative
A common RCA failure mode is collapsing timeline-building and cause-finding into a single step — the team writes "radiologist was too busy to call" directly onto the timeline, when in fact all that is objectively known is "critical result finalized at 14:32; verbal read-back documented at 19:50." The former is a hypothesis; the latter is a fact. Keeping these separate matters enormously, because premature causal framing anchors the whole team on one narrative (often the most visible or most recently discussed one) before the fishbone exercise has even begun — precisely the anchoring bias the RCA is trying to detect in the original clinical encounter.
Experienced RCA facilitators use a simple two-column format: the left column holds only verifiable timestamped facts, and the right column — populated only in the next stage — holds candidate contributing factors linked back to specific timeline entries. This structural discipline is one of the most effective and least expensive quality improvements a review program can adopt.
Contributing Factor Mapping — Building the Ishikawa Diagnostic-Error Fishbone
With the factual timeline in hand, the team now populates a fishbone (Ishikawa/cause-and-effect) diagram: a horizontal "spine" runs to the error outcome at the head, and each timeline gap or anomaly is sorted onto one of several categorical "bones." For diagnostic error specifically, the four bones that consistently capture the overwhelming majority of contributing factors are cognitive, system/process, communication, and patient-related.
- ~75%: Errors with ≥1 cognitive factor (illustrative, malpractice case-review studies)
- ~65%: Errors with ≥1 system factor (illustrative, same case-review studies)
- ~55–75%: Errors involving >1 category (multifactorial causation is the norm)
- ~7: Median contributing factors/case (illustrative RCA case-series average)
The four bones of a diagnostic-error fishbone
Cognitive factors capture failures in the clinician's reasoning process itself: anchoring on an initial impression and failing to revise it as new data arrives, premature closure (stopping the workup once a plausible-enough diagnosis is reached), availability bias (over-weighting a recent or memorable similar case), confirmation bias (selectively seeking data that supports the leading hypothesis), and framing effects from how the case was handed off or referred.
System/process factors capture the environment the clinician was reasoning inside: no closed-loop system for critical result notification, high census and cognitive load, fragmented records across sites, inadequate decision-support tooling, and absent forcing functions that would make the unsafe path harder to take than the safe one.
Communication factors capture breakdowns in the transfer of information between people: verbal-only handoffs without written confirmation, ambiguous or jargon-heavy discharge instructions, a critical value called but not read back or documented, and language or interpreter gaps at any touchpoint.
Patient-related factors capture legitimate clinical complexity that made the correct diagnosis genuinely harder to reach: atypical presentation, significant comorbidity masking red-flag findings, and structural barriers (transportation, insurance, health literacy) that delayed follow-up.
Crucially, the four-bone structure is a lens for organizing evidence, not a verdict — most confirmed diagnostic errors involve factors from at least two, often three, of these categories acting together, which is the core empirical justification for structured, multifactorial RCA over single-cause blame assignment.
The most common analytic error at this stage is stopping once a single satisfying explanation is found on one bone — for example, blaming a busy shift alone (system) while ignoring a documented anchoring bias (cognitive) that occurred earlier in the same case. A complete fishbone almost always has entries on three or four bones, not one.
From timeline gap to fishbone entry
Each factual gap identified in Stage 2 is examined by the team and mapped to one or more bones with a short evidentiary note. For instance, the "critical finding flagged at 14:32, acknowledged at 19:50" gap might map to two separate entries: a system entry ("no auto-escalation policy for unacknowledged critical PACS flags after 60 minutes") and a communication entry ("verbal critical-result call not documented per policy"). Splitting a single gap into multiple category entries when multiple plausible mechanisms exist is standard practice and is what the "Analysis Depth" control in this simulation represents — deeper analysis surfaces more of these compound, cross-category entries rather than accepting the first, simplest explanation.
Facilitation technique — silent brainstorming before category sorting
A well-run fishbone session begins with a silent, individual brainstorming round: every team member — the ED physician, the radiologist, the nurse involved, a patient-safety officer, and where feasible a non-clinical human-factors specialist — writes candidate contributing factors on cards without discussion. Only after all cards are collected are they read aloud, clustered by theme, and sorted onto the four bones as a group.
This silent-first structure exists specifically to counter groupthink and hierarchy effects: if the attending physician speaks first and proposes "the ED was just too busy that night," more junior team members are measurably less likely to independently raise a cognitive-bias explanation that might implicate the physician's own reasoning, even if they privately believe it is the more important factor. Written, anonymous-if-desired brainstorming followed by group sorting has been shown in patient-safety facilitation literature to surface a broader, more balanced set of contributing factors across all four categories than open floor discussion alone.
Applying the 5 Whys — Separating Contributing Factors from True Root Causes
Not every entry on the fishbone is a root cause. A contributing factor is anything that plausibly increased the likelihood of the error; a root cause is the deepest fixable point in a causal chain — the one where a systemic change would have prevented not just this error but a whole class of similar future errors. The "5 Whys" technique is the standard tool for walking each bone down to that level.
- 3–6: Typical "Whys" to reach a root cause (illustrative, varies by causal chain)
- ~3:1: Contributing factors → true root causes (illustrative convergence ratio)
- <20%: Cases with a single dominant root cause (multifactorial causation predominates)
- ~75–90%: Preventability score after full RCA (illustrative, structured-RCA cohorts)
The 5 Whys, applied to a diagnostic error
Consider one entry on the communication bone: "critical value called but not documented." Why #1: why wasn't it documented? Because there is no mandatory structured field for critical-value read-back in the nursing note template. Why #2: why is there no such field? Because the EHR order-entry team was never asked to add one when the critical-value policy was last updated. Why #3: why weren't they asked? Because policy updates and EHR build requests are owned by two different committees with no standing liaison between them. Why #4: why is there no liaison? Because the hospital's patient-safety governance structure was never redesigned after the two departments merged eighteen months earlier.
The chain stops at Why #4 because it has reached a systemic, fixable condition — a governance gap — rather than an individual's momentary lapse. Note that "the nurse forgot to document it" is never an acceptable stopping point in a mature RCA: individual lapses are treated as expected, inevitable noise in any human system, and the analysis must continue until it reaches the system-level condition that made the lapse likely or consequential.
Preventability scoring and convergence across bones
As causal chains from different bones are walked down, they frequently converge on a small number of shared root causes — a single governance gap, for instance, might be the terminus of both a communication-bone chain and a system-bone chain. This convergence is a good sign: it means the team is finding leverage points rather than an unmanageable list of unrelated fixes. Preventability is then scored, often on a structured ordinal scale (e.g., the Institute for Healthcare Improvement / NCC MERP-derived preventability frameworks used in many malpractice and quality-review programs), reflecting how confidently the panel believes the error would not have occurred had the identified root causes been absent. Scores in the 75–90% range are common for cases that reach a well-supported, multifactorial root-cause set; lower scores usually indicate either an irreducibly difficult clinical presentation or an incomplete analysis that has not yet reached true systemic causes.
A root cause has two defining tests: (1) if it were fixed, would this class of error become meaningfully less likely across many future patients, not just this one? and (2) is it within the organization's power to change? A finding that fails either test — such as "the disease has an atypical presentation" — is a contributing factor to document, not a root cause to act on.
Corrective Action Plan & Systems Fix — Closing the Loop
The final output of an RCA is only as good as the corrective actions it produces. A long list of well-identified root causes with no durable fix is a wasted 45 days. The strongest RCA programs rank candidate interventions using a hierarchy of effectiveness borrowed from human-factors engineering, favoring changes that make the unsafe path structurally harder to take over changes that merely ask people to remember to be more careful.
- ~40–60%: Actions rated "stronger" (system-level) (target mix in mature RCA programs)
- <20%: Actions rated "weaker" (training/reminder) (target ceiling; overused historically)
- 90 days: Re-audit interval for closed actions (typical follow-up checkpoint)
- ~30–50%: Recurrence reduction, closed-loop tools (illustrative, closed-loop alert literature)
The hierarchy of corrective-action effectiveness
From strongest to weakest: forcing functions and constraints (the unsafe action becomes physically or digitally impossible — e.g., an order cannot be finalized until a critical-result acknowledgment field is completed); automation and computerization (closed-loop critical-result tracking that auto-escalates to a supervisor if unacknowledged within 60 minutes); simplification and standardization (a single structured handoff template replacing ad hoc verbal sign-out); redundancies and double-checks (mandatory independent second read of any imaging study ordered from the ED for a high-risk chief complaint); and, weakest but still sometimes necessary, rules, policies, training, and reminders (an updated critical-values policy, a refresher module on cognitive bias).
Historically, RCA action plans over-relied on the weakest tier — a new policy memo or a mandatory training module — because these are cheap and fast to implement and look responsive to regulators. But policies and training alone have consistently weak, poorly sustained effects on error recurrence, because they depend on every individual remembering and complying every time under exactly the conditions (fatigue, time pressure, interruption) that produced the original error. Mature RCA governance now requires that at least some meaningful share of the action plan sit in the stronger tiers — forcing functions or automation — for the plan to be considered adequate.
A useful test before finalizing any corrective action: "if the exact same root cause condition recurred tomorrow, would this action actually prevent the error, or would it merely make someone feel that something was done?" Actions that only pass the second bar should be paired with, not substituted for, a stronger-tier fix.
Ownership, timelines, and measuring whether the fix worked
Each corrective action is assigned a named owner (not a department), an implementation deadline, and — critically — a measurable outcome metric that will be checked at a defined re-audit interval, commonly 90 days. For a closed-loop critical-result alert system, the outcome metric might be "percentage of critical PACS flags acknowledged within 60 minutes," tracked before and after implementation. If the metric does not move, the action is judged to have failed regardless of whether it was technically "implemented," and the RCA is reopened.
The final report is also fed back into the organization's aggregate patient-safety data: individual RCAs are periodically pooled and re-analyzed for recurring root-cause themes across many unrelated cases (a "meta-RCA"), which is often how a hospital discovers that the same governance gap identified in one case is quietly present behind several other, superficially unrelated safety events — turning a single diagnostic error into the trigger for a system-wide fix.
Sharing the finding without repeating the harm
The final step of a mature diagnostic-error RCA is disclosure and dissemination, handled as two related but distinct obligations. First, transparent communication with the patient and family about what happened, why, and what is being done to prevent recurrence — increasingly recognized as both an ethical obligation and, per communication-and-resolution program research, associated with lower rates of subsequent litigation compared with a defensive, disclosure-averse posture. Second, de-identified dissemination of the causal findings and corrective actions to the broader clinical department or system, so that the lessons generalize beyond the single unit where the error occurred.
Programs that close this loop well typically publish a brief, standardized case summary (chronology, contributing factors by category, root causes, corrective actions, and outcome metric) to a shared morbidity-and-mortality or patient-safety forum, explicitly stripped of any language that assigns individual blame. Over time, aggregating these summaries across many cases is what allows a health system to move from reactively fixing one error at a time to proactively redesigning the diagnostic process itself.
This simulation focuses on root cause analysis of diagnostic errors. It helps users identify the underlying factors that contribute to misdiagnosis and develop strategies to prevent such errors in clinical practice.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install