HomeRare Disease Diagnostic OdysseyUndiagnosed Disease Network Case Matching

🧭 Undiagnosed Disease Network Case Matching

Matching cases in the undiagnosed disease network to find similar patients for research and care.

Rare Disease Diagnostic Odyssey2DModerate60 FPS
undiagnosed-disease-network-case-matching ↗ Open standalone

Case Submission — Building a Shared Network Out of Isolated "N of 1" Patients

Every year, thousands of patients complete an exhausting "diagnostic odyssey" — years of testing, specialists, and inconclusive genome sequencing — and land on a single unexplained candidate variant in a gene of unknown significance. Alone, that finding is unprovable: one patient with one variant is a coincidence, not evidence. The Undiagnosed Diseases Network (UDN) and sister platforms exist to solve exactly this problem, by making every submitted case discoverable by every other clinician and researcher on Earth.

  • 2014: UDN founded (NIH Common Fund initiative)
  • 12+: UDN clinical sites (academic medical centers, US)
  • >2,000: Cases evaluated to date (accepted for full UDN work-up)
  • ~19: Avg. prior tests per patient (before UDN acceptance)

What gets submitted into the network

A UDN or GeneMatcher case submission is a structured record, not a free-text chart note. The core elements are:

• Candidate gene(s)/variant(s): typically from exome or genome sequencing — a variant of uncertain significance (VUS) in a gene with little or no prior disease association • Structured phenotype: clinical features coded with standardized terms (see Stage 3) rather than prose, so they are machine-comparable across languages and institutions • Inheritance pattern and family structure: trio/quad sequencing data, segregation within the pedigree • De-identified metadata: age of onset, ethnicity, prior test history — enough to judge similarity, not enough to identify the person

Because the record is structured, a case submitted in Boston can be automatically compared against one submitted in Tokyo or Melbourne without a human ever manually reading both charts first.

A single undiagnosed patient with a VUS in a poorly characterized gene is statistically almost impossible to publish as a new disease. The entire matchmaking model exists to convert that unpublishable "n of 1" into a defensible "n of 2 or more" — see Stage 4.

Why a single network node, not isolated hospital databases

Before matchmaking platforms existed, a candidate gene finding usually stayed trapped in the lab notebook or EHR of the clinician who found it. Two labs on opposite sides of the world could independently sequence patients with damaging variants in the very same gene and never learn of each other — the evidence needed to confirm causality simply never met.

The UDN and its international counterparts (Undiagnosed Diseases Network International, UDNI) solve this by pooling submissions into common, queryable infrastructure connected through Matchmaker Exchange (Stage 2), so a match can surface automatically the moment a second matching case is entered anywhere in the federation — sometimes years after the first submission.

GeneMatcher & Matchmaker Exchange — Searching the World for the Same Broken Gene

Matchmaker Exchange is a federated search protocol connecting independent databases — GeneMatcher, PhenomeCentral, DECIPHER, MyGene2, and others — so a query submitted to any one of them searches all of them simultaneously. Instead of one giant central database, each institution keeps control of its own patient data while still exposing it to a shared search layer, letting a researcher in one country discover a matching gene finding from a lab that never would have otherwise crossed their path.

  • 2013: GeneMatcher launched (Baylor College of Medicine)
  • >15,000: GeneMatcher submitters (researchers & clinicians, 90+ countries)
  • 10+: Matchmaker Exchange nodes (federated databases searched at once)
  • >14,000: Genes submitted to GeneMatcher (candidate genes of interest)

How the search actually works

Each node exposes a simple, privacy-preserving API: submit a gene symbol (and optionally variant-level detail), get back a list of other submitters who have flagged the same gene — no raw patient data changes hands at this stage, only contact metadata for the submitters. Matching is continuous: a query left standing today can surface a match a year from now, the moment someone else in the world submits the same gene.

Matching signal comes in tiers of strength: • Exact same variant in the same gene (strongest) • Different variants, same gene, overlapping predicted loss-of-function mechanism • Same gene, different but biologically plausible variant class (e.g., missense clustering in the same protein domain) • Same pathway/protein complex, different gene — weaker signal, used to nominate new candidate genes in known disease mechanisms

Statistical power of even one additional match

Rare disease genetics is a probability argument. Any human genome carries roughly 1–2 rare, potentially damaging variants in genes of unknown significance purely by chance — so one patient with one variant proves nothing. But the probability that two unrelated patients, evaluated independently on different continents, would both carry a rare damaging variant in the exact same gene AND share overlapping, specific clinical features is vanishingly small unless that gene truly causes the phenotype.

This is why matchmaking is so disproportionately powerful: going from n=1 to n=2 does not double the evidence, it can move a finding from "unpublishable" to "highly likely causal" in a single step, especially once phenotype similarity (Stage 3) is layered on top.

Analyses of solved UDN and GeneMatcher cases show that a large share of newly confirmed disease genes were established with as few as 2–5 matched, unrelated families worldwide — evidence that would have been invisible without cross-institutional search.

The Human Phenotype Ontology — Scoring Clinical Similarity Between Candidate Matches

A shared gene is only half the argument. The other half is whether two patients actually look clinically alike. The Human Phenotype Ontology (HPO) gives clinicians a standardized, hierarchical vocabulary — over 18,000 terms — for coding clinical abnormalities, so "unusually small head" and "microcephaly" become the identical, computable term regardless of who wrote the note or in what language, allowing algorithms to compute a real similarity score between two patients instead of a subjective impression.

  • 18,000+: HPO terms (standardized clinical abnormalities)
  • 2008: HPO first released (now maintained by global consortium)
  • Exomiser, Phenomizer: Phenotype-matching tools (semantic similarity scoring)
  • 10–30: Typical HPO terms/case (coded per UDN patient)

How similarity is actually computed

HPO terms sit in a directed acyclic graph — a hierarchy from broad ("Abnormality of the nervous system") down to highly specific ("Episodic ataxia"). Two patients rarely share the exact same fine-grained term, but semantic similarity algorithms (e.g., Phenomizer, Exomiser's hiPHIVE) measure how closely related their terms are in the ontology graph, then combine per-term similarity into a single case-level score.

The score weights rare, specific terms far more heavily than common ones: two patients both having "developmental delay" is weak evidence (it appears in thousands of conditions); two patients both having "episodic ataxia" plus "paroxysmal dystonia" plus "onset in infancy" is strong evidence, because that specific combination is rare in the general undiagnosed population.

Combining gene evidence and phenotype evidence

The real diagnostic power comes from multiplying the two independent signals from Stages 2 and 3 together: variant-level match probability × phenotype similarity score. A weak gene signal (e.g., same pathway, different gene) paired with an extremely strong, specific phenotype overlap can still justify a comparison; a strong gene signal paired with divergent phenotypes may indicate variable expressivity of the same condition rather than a false match.

Most matchmaking pipelines require candidate pairs to clear a minimum combined threshold before a connection is even surfaced to clinicians — filtering out the large number of coincidental, low-value matches that would otherwise overwhelm a manually reviewed queue.

Standardized phenotyping is what makes automated, cross-language matching possible at all — without it, "small head" (English), "microcéphalie" (French), and "小头畸形" (Chinese) would never be recognized as the same clinical finding by a search algorithm.

Match Confirmed — Connecting Clinical Teams Across Institutions and Borders

A high-confidence computational match is only the beginning. Matchmaker Exchange and GeneMatcher deliberately stop short of sharing full patient data automatically — instead, they connect the submitting researchers and clinicians directly, by email or secure message, so the humans can review case details, confirm the match is genuine, obtain appropriate consent, and — if warranted — pool sequencing data, imaging, and clinical notes for formal case-series analysis.

  • Days–years: Typical time to first match (depends on gene rarity)
  • ~35%: UDN diagnostic rate (of previously unsolved cases accepted)
  • Hundreds: Cases solved via matchmaking (documented disease-gene discoveries)
  • Contact info only: Data shared at match time (full data shared after consent)

From automated match to human confirmation

When GeneMatcher or a UDN clinical site flags a candidate match, the system sends both submitters an introduction — typically nothing more sensitive than "you and another submitter both have an interest in gene X; here is their contact information." From there:

1. The two teams exchange de-identified case summaries to sanity-check the match 2. If phenotype and variant evidence still look compelling, they pursue data sharing agreements appropriate to their institutions and jurisdictions 3. Segregation analysis, functional studies (model organisms, cell assays), and additional family recruitment follow to build formal causality evidence (following frameworks like ClinGen's gene-disease validity criteria) 4. A case series manuscript is prepared once enough independent families are assembled — often 3 or more unrelated probands with convergent functional evidence

Why the human-in-the-loop step matters

Automated matching surfaces candidates; it does not adjudicate them. Two patients can share a rare variant in the same gene by chance, or a phenotype overlap that reflects a different, more common condition. Direct clinician-to-clinician contact catches these false positives quickly, while also uncovering subtler true positives that a similarity score alone might rank lower — an experienced clinician may recognize a shared, unusual facial gestalt or disease trajectory that phenotype coding did not fully capture.

International Undiagnosed Diseases Network (UDNI) partners now span more than 20 countries, meaning a match confirmed today may connect a family in rural Australia with a research lab in Baltimore that has spent a decade studying the exact same orphan gene.

Novel Disease-Gene Discovery — From Matched Cohort to a Diagnosis the Next Patient Can Receive

Once enough independent, matched families converge on the same gene with overlapping phenotypes and supporting functional evidence, the finding graduates from "candidate" to a formally described, published disease-gene relationship — entered into clinical databases such as OMIM and ClinVar, and immediately becomes searchable by every diagnostic lab and clinician worldwide. The undiagnosed patient who started the search becomes patient zero of a named condition, and every subsequent patient with the same gene now gets a diagnosis in days instead of years.

  • Hundreds: New disease-gene relationships/yr (via international matchmaking)
  • >2,500: ClinGen gene curations (formally assessed gene-disease pairs)
  • 3–8: Median cohort size at publication (unrelated matched families)
  • Days: Diagnostic turnaround after publication (vs. years for the original cohort)

What "confirmed" requires

Establishing a new gene-disease relationship generally follows a structured evidence framework (e.g., ClinGen Gene-Disease Validity criteria), combining:

• Genetic evidence: multiple unrelated probands with rare, predicted-damaging variants in the same gene; segregation with disease in families where possible • Variant-type consistency: a plausible shared mechanism (e.g., all loss-of-function, or missense variants clustering in the same functional domain) • Functional evidence: model organism (zebrafish, Drosophila, mouse) or cell-based assays showing the variant disrupts gene function in a way that explains the phenotype • Phenotypic consistency: the matched cohort shares a recognizable, specific clinical picture distinguishable from other known conditions

The compounding effect on future patients

Every confirmed discovery immediately reduces the odds of future patients having to endure their own multi-year diagnostic odyssey — the moment a gene-disease pair is published and entered into ClinVar/OMIM, standard diagnostic exome and genome pipelines everywhere in the world will flag it automatically. The network effect compounds: as more sites join, more genomes are searchable, more matches surface, and the pool of candidate genes still needing that first confirming match keeps shrinking.

Well-known UDN and international matchmaking success stories include the description of NGLY1 deficiency and multiple KIF1A-, TANGO2-, and WDR45-related disorders — each first identified as an unexplained single case, then confirmed once matchmaking surfaced additional unrelated patients with the same gene and overlapping phenotypes.
⚙ Under the hood

Matching cases in the undiagnosed disease network to find similar patients for research and care.

CanvasBiomedicine

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)