Assembling a Diagnostic-Concordance Case Set
A representative case panel anchors any fair AI diagnostic benchmark.
- 3: Case categories tracked (common, moderate, rare/atypical)
- 8: Diagnoses per category pool (condition label options)
- Physician: Ground truth source (confirmed final diagnosis)
- 10–200: Panel size range (cases per benchmark run)
Why case-mix matters
A skewed panel toward easy cases inflates apparent accuracy.
Physician-confirmed labels
Each case carries a real physician diagnosis as ground truth.
The AI Triage Bot Suggests a Diagnosis
For every case, the bot outputs one likely underlying condition.
- 1 dx: Inference target (top suggested diagnosis)
- Symptoms: Input signal (patient-reported intake data)
- Label: Output form (single condition name)
- Broad: Scope (not just emergency triage)
Suggestion, not confirmation
The bot proposes a diagnosis without ordering confirmatory tests.
General diagnostic scope
This benchmark spans many conditions, not one binary decision.
Comparing Bot Suggestion to Physician Diagnosis
Each bot label is checked against the matching physician label.
- 1 case: Comparison unit (bot label vs MD label)
- Green: Match outcome (diagnoses agree)
- Red: Mismatch outcome (diagnoses disagree)
- Exact dx: Granularity (not just triage tier)
Exact-match scoring
A match requires the same named condition, not just a category.
Case-by-case visibility
Every case is inspectable, not just the aggregate score.
Calculating the Overall Concordance Rate
Matches divided by total cases gives one headline accuracy number.
- Concordance %: Metric (bot-physician agreement rate)
- ~88%: Typical MD-MD agreement (inter-rater benchmark)
- Stabilizes: Sample size effect (with more cases tested)
- Hides spread: Single number risk (across case difficulty)
One headline number
Overall concordance summarizes performance across the whole panel.
Physician benchmark anchor
Bot rate is compared against typical physician-physician agreement.
Accuracy Benchmark By Case Type
Concordance is split by common, moderate, and rare case categories.
- High: Common-case concordance (bot closely tracks physicians)
- Mid: Moderate-case concordance (noticeable accuracy drop)
- Low: Rare-case concordance (largest bot-physician gap)
- Targeting: Use of breakdown (guides model improvement)
Bar chart by category
Concordance bars reveal exactly where the bot struggles most.
Guiding model improvement
Category gaps point developers toward the weakest condition types.