HomeVoice & Swallowing Disorder DiagnosticsVoice Therapy Outcome Acoustic Analysis Simulator

🗣 Voice Therapy Outcome Acoustic Analysis Simulator

This simulation allows users to analyze the acoustic outcomes of voice therapy treatments. It provides detailed information on changes in pitch, intensity, and quality of the voice before and after therapy sessions.

Voice & Swallowing Disorder Diagnostics2DModerate60 FPS
voice-therapy-acoustic-analysis ↗ Open standalone

Baseline Dysphonic Voice — Acoustic Recording

Before any intervention begins, clinical voice evaluation starts with a standardized acoustic recording: a sustained vowel /a/ held for 3–5 seconds plus a connected-speech passage (e.g. the Rainbow Passage), captured at 44.1 kHz in a low-noise environment. This recording is the objective anchor against which every subsequent therapy session will be measured.

  • <1.04%: Normal jitter (local) threshold (Praat/MDVP cutoff)
  • <3.81%: Normal shimmer (local) threshold (Praat/MDVP cutoff)
  • 2–4 dB: Dysphonic baseline CPPS (sustained vowel, severe cases)
  • 61–120: Severe VHI range (of 120-point scale)

What acoustic analysis is measuring

Dysphonia — an umbrella term for abnormal voice quality — arises from irregular vocal fold vibration. Causes range from vocal fold nodules and polyps (phonotrauma from chronic hyperfunctional voice use), to muscle tension dysphonia, to neurological conditions like spasmodic dysphonia or vocal fold paresis. Regardless of etiology, the perceptual impression of a "rough," "breathy," or "strained" voice corresponds to measurable irregularities in the glottal cycle.

A healthy vocal fold vibrates with near-periodic regularity — each glottal cycle closely resembles the one before it in both duration and amplitude. In dysphonia, that regularity breaks down: successive cycles vary in length (frequency perturbation) and in amplitude (intensity perturbation), and turbulent airflow through an incompletely closing glottis adds aperiodic noise energy across the spectrum.

Acoustic analysis converts a subjective complaint ("my voice sounds rough") into a reproducible number. This objectivity is what allows a clinician to demonstrate — not just assert — that therapy is working.

Recording protocol matters as much as the algorithm

Acoustic measures are highly sensitive to recording conditions, and clinical guidelines (e.g. ELS, ASHA) specify strict protocols to keep results comparable across sessions:

• Microphone: head-mounted condenser microphone at a fixed 4–10 cm mouth-to-mic distance, avoiding handheld variability • Environment: ambient noise below 50 dB SPL, non-reverberant room • Task: sustained /a/ or /i/ at comfortable pitch and loudness, held ≥3 seconds after onset transients are excluded, plus a standardized connected-speech passage • Sampling: 44.1 kHz / 16-bit minimum, since jitter/shimmer algorithms are sensitive to sampling artifacts at lower rates

Without this consistency, session-to-session changes in a jitter or CPPS value could reflect microphone placement drift rather than true vocal fold improvement — a critical confound when acoustic measures are used to track therapy outcomes over months.

Who these patients are

Voice disorders affect an estimated 1 in 13 adults in a given year, and lifetime prevalence in occupational voice users (teachers, singers, clergy, call-center staff) can exceed 50%. Common referral diagnoses for acoustically-monitored voice therapy include:

• Muscle tension dysphonia (MTD) — hyperfunctional laryngeal posturing without structural lesion • Vocal fold nodules/polyps — phonotraumatic lesions from repeated forceful vocal fold collision • Presbyphonia — age-related vocal fold atrophy and bowing • Unilateral vocal fold paresis — reduced closure from impaired innervation

Acoustic baseline recording establishes disorder severity, screens for red flags requiring laryngoscopy or ENT referral, and — most importantly for this simulation — sets the numerical starting point for tracking recovery.

Spectrographic Analysis — Jitter, Shimmer & CPPS

The recorded waveform is run through acoustic analysis software (Praat, MDVP, VoxMetria) to extract quantitative perturbation measures. Classic time-domain measures — jitter and shimmer — require the software to first find each individual glottal cycle, which becomes unreliable exactly when the voice is most disordered. Cepstral analysis solves this by working in the frequency domain instead.

  • 50–150: Typical analysis window (glottal cycles per sample)
  • r ≈ 0.7–0.8: CPPS–perceptual severity correlation (vs GRBAS/CAPE-V ratings)
  • 2–6%: Jitter in severe dysphonia (vs <1.04% normal)
  • 8–18%: Shimmer in severe dysphonia (vs <3.81% normal)

Jitter and shimmer — and why they fail in severe dysphonia

Jitter (local) quantifies cycle-to-cycle variation in fundamental period:

Jitter(%) = [ (1/(N−1)) Σ|Tᵢ − Tᵢ₊₁| ] / [ (1/N) Σ Tᵢ ] × 100

Shimmer (local) quantifies cycle-to-cycle variation in peak amplitude using the analogous formula on cycle amplitudes Aᵢ instead of periods Tᵢ.

Both measures depend entirely on the algorithm correctly identifying each glottal cycle boundary (pitch-period detection). In mild-to-moderate dysphonia this works reasonably well. But in moderate-to-severe dysphonia — and especially in voices with significant aperiodicity or breathiness — the software frequently mis-detects cycle boundaries, and jitter/shimmer values become unreliable or simply cannot be computed at all. This is precisely the population for whom clinicians most need a robust measure.

CPPS — a period-independent alternative

Cepstral Peak Prominence Smoothed (CPPS) sidesteps cycle-detection entirely. The cepstrum is computed by taking the power spectrum of the log power spectrum of the signal — a "spectrum of a spectrum." A periodic voice signal produces a sharp, prominent peak in the cepstrum at the quefrency corresponding to the fundamental period; an aperiodic, noisy voice produces a flattened, less prominent peak.

CPPS measures the amplitude of that cepstral peak relative to a linear regression line fit through the overall cepstrum, after smoothing across both time and quefrency. Because it does not require identifying individual glottal cycles, CPPS remains computable and clinically meaningful even in severely dysphonic, near-aperiodic voices where jitter and shimmer break down — making it the current gold-standard acoustic measure recommended by the American Speech-Language-Hearing Association (ASHA) expert panel for tracking dysphonia severity and treatment outcomes.

CPPS values below roughly 4.1 dB on sustained vowels are strongly associated with perceptually severe dysphonia in multiple validation studies, while CPPS above ~8–14 dB (task-dependent) is typical of normophonic voices — giving clinicians a continuous severity scale that survives even when jitter/shimmer cannot be extracted.

Reading the spectrogram directly

Alongside numeric measures, clinicians visually inspect a narrowband spectrogram of the sample. In a healthy voice, harmonics appear as thin, sharply-defined, evenly-spaced horizontal bands stacked at integer multiples of F0, with a low noise floor between them.

In a dysphonic voice, several visual signatures appear: harmonic bands widen and blur (frequency instability), the noise floor between harmonics rises and can fill in entirely at higher frequencies (turbulent, breathy airflow), and sub-harmonic or aperiodic energy can appear as diffuse, disorganized coloring rather than clean lines. Spectrographic pattern, jitter/shimmer, and CPPS together form a converging, multi-method picture of vocal fold vibratory behavior that guides both diagnosis and, in this simulation, the visual "cleanup" of the signal as therapy progresses.

Semi-Occluded Vocal Tract (SOVT) — Straw Phonation Therapy

The therapeutic core of most modern voice rehabilitation programs is the semi-occluded vocal tract (SOVT) exercise: phonating through a narrowed opening — a thin straw, lip trill, or /u/ vowel — which raises supraglottal (oral) pressure and reflexively reduces the force with which the vocal folds collide, all without the patient needing to consciously relax the larynx.

  • 5–10: Intraoral pressure, straw phonation (cmH₂O vs ~2–4 normal /a/)
  • ~15–30%: Vocal fold collision force reduction (estimated, SOVT vs open vowel)
  • 5–28 cm: Typical clinical straw (length, ~5 mm bore)
  • 5–10 min: Home exercise dose (sets, 2–3×/day)

The physics of back-pressure and vocal fold sparing

When a narrow tube is added at the lips, airflow exiting the vocal tract is restricted, and air pressure builds up in the mouth and pharynx just above the glottis. This supraglottal pressure partially counteracts the pressure driving the vocal folds apart during the open phase of each vibratory cycle, which — through the aerodynamic-myoelastic mechanics of phonation — reduces the peak velocity and force with which the folds slam back together during closure.

Lower collision force means lower mechanical impact stress on vocal fold tissue per cycle. For patients with nodules, polyps, or general phonotraumatic hyperfunction, this is the direct therapeutic mechanism: SOVT exercises let the patient continue phonating (rather than resting the voice entirely) while measurably reducing the tissue trauma that caused or maintains the lesion.

Vocal tract inertance and immediate vibratory stabilization

Beyond reducing collision force, SOVT exercises increase the acoustic inertance of the vocal tract — its opposition to rapid changes in airflow. Higher inertance improves impedance matching between the vocal tract and the glottal source, which has two visible acoustic effects even within a single exercise: vocal fold vibration becomes more symmetric and periodic in real time, and phonation requires measurably less subglottal pressure and laryngeal muscular effort to sustain — i.e. it becomes easier to produce voice with less "push."

This is why, in the simulation, the waveform visibly stabilizes the moment straw phonation begins — the effect is not a slow retraining process but an immediate biomechanical consequence of the semi-occlusion itself, layered on top of the slower, cumulative retraining that occurs across the full therapy course.

SOVT exercises are unusual among rehabilitative techniques in that they produce an immediate, within-trial acoustic improvement (measurable in seconds) as well as a cumulative, session-over-session training effect (measurable in weeks) — which is why they anchor nearly every evidence-based voice therapy protocol.

Comparing the major evidence-based voice therapy approaches

SOVT/straw phonation is rarely used alone — it is typically one component within a broader, individualized therapy plan drawn from several evidence-based approaches, each targeting slightly different aspects of phonatory physiology and each with its own evidence base and target population, summarized below.

Evidence-based voice therapy technique comparison

ProductIndicationTrial DesignKey Result
SOVT / Straw PhonationPhonotrauma, nodules/polyps, MTD, general hyperfunctionSemi-occlusion raises supraglottal pressure, lowers vocal fold collision force, improves source-tract impedance matchImmediate within-trial acoustic effect; easy home carryover
Resonant Voice Therapy (RVT)MTD, phonotraumatic lesions, vocal fatigueTrains forward oral vibratory sensations (e.g. humming into /m/-vowel glides) at the barely-abducted vocal fold configuration that maximizes vocal economyStrong efficacy evidence; low phonatory effort
Vocal Function Exercises (VFE)Vocal fold atrophy, presbyphonia, general strengtheningSystematic warm-up/stretch/contract/adduction exercise set performed like physical therapy for the laryngeal musculatureRebuilds vocal fold muscle tone and stamina over a full course
LSVT LOUDParkinson's disease and other hypokinetic dysarthriasIntensive recalibration of vocal loudness/effort perception via high-effort daily practice over 4 weeksOnly voice therapy with Level I evidence in Parkinson's disease

Progressive Acoustic Normalization Across Sessions

A single therapy session rarely resolves chronic dysphonia. Voice therapy is a rehabilitative course, typically 6–12 sessions delivered weekly or biweekly over 6–8 weeks, layered on a daily home exercise program. Acoustic measures — recorded at the start of every clinic visit under the same protocol as baseline — provide the objective trend line that guides whether to continue, intensify, or discharge therapy.

  • 6–12: Typical therapy course length (sessions over 6–8 weeks)
  • +1–3 dB: CPPS gain by mid-course (typical by session 5–6)
  • ~70–80%: Patients showing measurable gain by session 6 (evidence-based programs)
  • ≈18 pts: VHI minimal clinically important difference (reduction threshold)

Why change is gradual, not immediate

While SOVT exercises produce an immediate within-trial acoustic effect, durable change requires motor retraining: the patient's habitual, largely unconscious laryngeal muscle activation patterns must be rebuilt through repetition until the new, lower-effort phonatory pattern becomes the default rather than something requiring conscious effort.

This retraining follows a classic motor-learning curve — larger gains early, with diminishing but still meaningful gains later, and some regression between sessions as old habitual patterns reassert themselves before the new pattern consolidates. Consequently, week-to-week acoustic measures do not fall in a straight line toward normal; clinicians expect a generally improving trend with normal session-to-session variability layered on top, and interpret the trend over 3–4 data points rather than any single visit in isolation.

What drives the between-session slope

The rate of acoustic normalization across a therapy course depends on several factors reproduced in this simulation's progression:

• Baseline severity and duration of hyperfunctional pattern — longer-standing hyperfunction takes longer to retrain • Home practice adherence — patients completing daily SOVT sets show faster acoustic normalization than clinic-only attendance • Presence and resolution of structural lesions — nodules/polyps may need several weeks of reduced collision force before tissue swelling subsides and the mucosal wave normalizes • Co-occurring factors — reflux, allergy, vocal load at work/home, and hydration all modulate the trajectory session to session

Clinically, a plateau in CPPS or VHI across 2–3 consecutive sessions is a common trigger to reassess the treatment plan, check adherence, or consider adjunct medical management (e.g. laryngopharyngeal reflux treatment) rather than simply continuing the same exercises indefinitely.

Because CPPS remains reliably computable throughout the entire severity range — unlike jitter and shimmer, which can fail to compute in the most severe early sessions — it is uniquely suited to plotting the full session-by-session recovery trajectory from baseline to discharge on one continuous scale.

Combining acoustic and patient-reported measures

Acoustic normalization is tracked alongside the Voice Handicap Index (VHI) — a 30-item patient-reported questionnaire covering functional, physical, and emotional impact of the voice disorder, scored 0–120 (mild 0–30, moderate 31–60, severe 61–120). VHI captures the patient's lived experience (e.g. difficulty being heard on the phone, social withdrawal, vocal fatigue) that a lab acoustic measure alone cannot.

The two measures usually — but not always — move together. A patient can show excellent acoustic normalization while still reporting handicap due to residual vocal fatigue in specific situations, or vice versa; discrepancies between the two are clinically informative and often prompt targeted counseling or vocal hygiene education alongside the exercise program.

Outcome — Normalized CPPS & Voice Handicap Index

At the end of a successful therapy course, the objective acoustic profile and the patient-reported outcome converge: CPPS has risen into or near the normative range, jitter and shimmer have fallen below pathological thresholds, and VHI has dropped by more than the minimal clinically important difference — together constituting evidence-based justification for discharge from active therapy.

  • ≥8–10 dB: Discharge CPPS target (sustained vowel/connected speech)
  • <1.0% / <3.8%: Discharge jitter/shimmer target (within normative bounds)
  • <10–18: Discharge VHI target (mild/resolved handicap)
  • 3–6 mo: Recommended follow-up check (post-discharge maintenance)

Confirming normalization, not just improvement

A statistically significant improvement in acoustic measures is encouraging but not, by itself, sufficient grounds for discharge — the goal is normalization (or the best achievable outcome given any residual structural pathology), not merely "better than before." Clinicians compare final-session values against published normative ranges: jitter under roughly 1.04%, shimmer under roughly 3.81%, and CPPS in the upper single digits to low double-digit dB range depending on task and analysis software.

When residual structural findings remain (e.g. a persistent small nodule visible on laryngoscopy), acoustic measures may plateau short of fully normative values even with excellent therapeutic response — in those cases, discharge criteria weigh functional and patient-reported outcome (VHI, return to occupational voice demands) alongside the acoustic profile rather than requiring numeric perfection.

The Voice Handicap Index as the outcome anchor

The VHI reduction is often the measure that matters most to the patient and to outcome research, since it directly quantifies restored quality of life. A drop of roughly 18 points or more is generally considered clinically meaningful (exceeding measurement noise and reflecting a real change the patient would notice), and many successful therapy courses move patients from the severe (61–120) or moderate (31–60) bands down into the mild (0–30) band.

Combining VHI with acoustic measures gives clinicians and researchers a dual-endpoint outcome model — an objective, instrument-derived measure (CPPS, jitter, shimmer) triangulated with a subjective, patient-derived measure (VHI) — which is now the standard reporting approach recommended in voice therapy outcome literature, since either measure alone can miss part of the clinical picture.

Objective acoustic confirmation of recovery — not just the patient's subjective sense of feeling better — is what allows voice therapy outcomes to be compared across clinics, published in outcome studies, and used to justify insurance coverage for continued or repeat therapy courses.

Maintenance and relapse prevention

Because the underlying behavioral pattern (vocal hyperfunction) can re-emerge under stress, illness, or high vocal load even after successful therapy, discharge typically includes a structured maintenance plan: continued light SOVT practice as a vocal "warm-up," vocal hygiene education (hydration, reflux management, avoiding throat clearing and shouting), and a scheduled follow-up acoustic check at 3–6 months to confirm the gains have held.

Relapse — a return of acoustic measures toward baseline — is itself detectable using the same protocol used throughout therapy, closing the loop: the objective measurement system that diagnosed the disorder, guided the therapy course, and confirmed the outcome also serves as the long-term surveillance tool for maintaining vocal health.

⚙ Under the hood

This simulation allows users to analyze the acoustic outcomes of voice therapy treatments. It provides detailed information on changes in pitch, intensity, and quality of the voice before and after therapy sessions.

CanvasBiomedicine

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)