🎙 Multilingual AI Scribe Accuracy Simulator
This simulation evaluates the accuracy of an AI scribe when documenting consultations in multiple languages, emphasizing the challenges and potential improvements needed for effective multilingual use.
The Visit Starts in the Patient's Language
Placeholder: ambient scribe listens in from the very first spoken word.
- 9: Languages modeled (placeholder short caption)
- 3: Resource tiers (placeholder short caption)
- Ambient audio: Input mode (placeholder short caption)
- 2: Downstream steps (placeholder short caption)
Placeholder section heading text
Placeholder body text, short and generic for now.
Speech-to-Text Accuracy Tracks Training Data Volume
Placeholder: word-level accuracy differs sharply by language resource level.
- ~97%: High-resource accuracy (placeholder short caption)
- ~88%: Medium-resource accuracy (placeholder short caption)
- ~76%: Low-resource accuracy (placeholder short caption)
- up to 18pt: Accent penalty (placeholder short caption)
Placeholder section heading text
Placeholder body text, short and generic for now.
Clinical Vocabulary Capture Lags Behind Plain Speech
Placeholder: medical terms are harder to capture than everyday words.
- ~95%: High-resource term acc. (placeholder short caption)
- ~82%: Medium-resource term acc. (placeholder short caption)
- ~65%: Low-resource term acc. (placeholder short caption)
- widens downstream: Gap vs. transcription (placeholder short caption)
Placeholder section heading text
Placeholder body text, short and generic for now.
Side-by-Side Accuracy Across Resource Tiers
Placeholder: the gap becomes visible once tiers sit side by side.
- ~96%: Best tier combined acc. (placeholder short caption)
- ~70%: Worst tier combined acc. (placeholder short caption)
- ~25pt: Tier spread (placeholder short caption)
- 9: Languages compared (placeholder short caption)
Placeholder section heading text
Placeholder body text, short and generic for now.
Accuracy Disparities Raise Real Equity Concerns
Placeholder: lower accuracy means more physician review is needed.
- ≥90% combined: Standard review threshold (placeholder short caption)
- <90% combined: Enhanced review threshold (placeholder short caption)
- Lower-resource speakers: Populations affected (placeholder short caption)
- Targeted physician review: Mitigation (placeholder short caption)
Placeholder section heading text
Placeholder body text, short and generic for now.
Placeholder highlight: equity requires monitoring accuracy per language.
This simulation evaluates the accuracy of an AI scribe when documenting consultations in multiple languages, emphasizing the challenges and potential improvements needed for effective multilingual use.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install