🎙 AI Ambient Scribe SOAP Note Generation Simulator
An AI-driven automatic generation of SOAP notes from an audio recording of a clinician-patient consultation.
Ambient Listening Begins With Explicit Consent
The visit starts like any other — consent shown, mic listening quietly.
- Always: Consent required (verbal or signed before capture)
- Passive audio: Capture mode (no manual note-taking needed)
- 10–20 min: Typical visit length (primary care average)
- Encrypted stream: Data handling (HIPAA-aligned transport)
Why consent comes first
Patients see a clear indicator before any audio is recorded.
Consent is logged alongside the encounter for audit purposes.
What the microphone hears
Ambient audio includes clinical talk plus ordinary small talk and noise.
Nothing is filtered yet — that happens in later stages.
Passive capture removes the need for manual typing during the visit.
Speech-to-Text Converts Audio Into Raw Transcript
A speech model turns the recorded conversation into timestamped text.
- ASR engine: Model type (automatic speech recognition)
- ~95%: Baseline accuracy (clean audio conditions)
- Diarized: Speaker labeling (doctor vs patient turns)
- Near real-time: Latency (streaming transcription)
From waveform to words
Audio frames convert into text tokens as the conversation unfolds.
Accuracy depends heavily on background noise and speaker overlap.
Handling imperfect audio
Noisy rooms and cross-talk lower transcription accuracy noticeably.
Speaker diarization helps separate doctor and patient utterances.
Cleaner audio input directly improves downstream note quality.
The LLM Filters Chatter And Extracts Clinical Content
A language model reads the transcript and pulls out relevant clinical facts.
- Raw transcript: Input (full conversation text)
- Non-clinical removed: Filtering (small talk discarded)
- 4: Output categories (Subjective, Objective, Assessment, Plan)
- Prompted LLM: Extraction method (structured section routing)
Separating signal from noise
Small talk about weather or parking gets dropped automatically.
Symptoms, exam findings, and plans get routed to their sections.
Mapping talk to structure
Each clinically relevant sentence is classified into S, O, A, or P.
The model reasons over context, not just keyword matching.
Structuring turns free-flowing conversation into a usable clinical record.
The Four SOAP Sections Populate With Drafted Text
Subjective, Objective, Assessment, and Plan fill with extracted content.
- Patient-reported: Subjective (symptoms and history)
- Exam findings: Objective (vitals and observations)
- Clinical impression: Assessment (diagnosis reasoning)
- Next steps: Plan (orders, meds, follow-up)
A familiar clinical format
SOAP structure mirrors how clinicians already organize their notes.
Each section fills incrementally as the LLM finishes drafting.
Draft, not final
The generated note is a starting point, not a finished record.
It still requires physician review before it becomes official.
A structured draft saves far more time than typing from scratch.
Physician Review, Edits, And Signs The Note
The doctor reviews the draft, corrects details, and signs off.
- Required: Review step (human-in-the-loop check)
- ~Minutes daily: Time saved (vs manual documentation)
- Light edits: Edit rate (typically minor corrections)
- Signed note: Final status (added to patient chart)
Why review still matters
The physician remains accountable for everything in the final note.
Quick edits catch any misheard or misclassified details.
The documentation payoff
Less typing during and after visits means more time with patients.
Ambient scribing aims to reduce clinician documentation burden.
Signed notes close the loop from conversation to chart in minutes.
An AI-driven automatic generation of SOAP notes from an audio recording of a clinician-patient consultation.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install