Language Barrier Identified
Doctor and patient share no common spoken language.
- 25.6M: US LEP patients (limited English proficiency)
- 2×: Adverse event risk (higher without interpreters)
- 350+: Languages spoken, US (across patient populations)
- ~15%: Encounters needing interpreters (of US clinical visits)
Why language gaps matter
Miscommunication drives misdiagnosis and unsafe care.
Professional interpretation cuts adverse events sharply.
Legacy interpretation methods
Phone lines and ad hoc family interpreters are slow and error-prone.
Where AI interpreters fit
On-device speech AI now sits directly inside the exam room.
Doctor Speaks — Live Capture & Transcription
Clinical speech is captured and transcribed as it happens.
- <5%: ASR word error rate (clinical speech models)
- ~300ms: Capture-to-text delay (streaming transcription)
- 100k+: Vocabulary size (medical terms indexed)
- 4: Mic array channels (noise-cancelling beamform)
Streaming speech recognition
Audio is chunked and transcribed continuously, not after silence.
Medical vocabulary tuning
ASR models are fine-tuned on clinical terminology and drug names.
Domain-tuned ASR sharply cuts term transcription errors.
Speaker diarization
The system tags who is speaking to route translation correctly.
AI Translation — Real-Time Language Conversion
Transcribed speech is translated into the patient's language instantly.
- 1-2s: Translation latency (end-to-end, streaming)
- 96%: Medical term accuracy (benchmark clinical corpora)
- 40+: Supported language pairs (in production systems)
- Full visit: Context window (conversation-aware model)
Neural machine translation
A transformer model converts meaning, not just words.
Clinical term protection
Drug names and dosages are locked against mistranslation.
Term-locking prevents dangerous dosage translation errors.
Latency vs accuracy tradeoff
Faster output can slightly reduce nuance and accuracy.
Patient Responds — Capture & Reverse Translation
Patient speech is captured and translated back to the doctor.
- 1.6s: Reverse latency (patient-to-doctor path)
- 92%: Accent robustness (across regional dialects)
- 95%: Symptom term recall (patient-reported terms)
- 98%: Turn detection accuracy (end-of-speech detection)
Symmetric interpretation path
The same pipeline runs in reverse for the patient's reply.
Handling accents and dialects
Models are trained across diverse regional speech patterns.
Broad accent training keeps accuracy stable across speakers.
Turn-taking detection
The system detects speech end to know when to translate.
Bidirectional Conversation Flow
Continuous real-time interpretation sustains full dialogue.
- 1.2s: Sustained latency (steady-state conversation)
- 6-20: Turns per encounter (typical consult length)
- 97%: Sustained accuracy (across full encounter)
- +34%: Patient satisfaction (vs phone interpretation)
Continuous cycling pipeline
Capture, translate, and speak repeat seamlessly each turn.
Conversation memory
Prior turns inform pronoun and context resolution.
Context memory keeps pronouns and follow-ups coherent.
Clinical outcomes
Full bidirectional flow restores natural doctor-patient rapport.