EMG-guided botulinum toxin therapy for laryngeal dystonia
Spasmodic dysphonia (laryngeal dystonia) is a focal, task-specific neurological voice disorder in which the intrinsic laryngeal muscles contract involuntarily during speech. Correctly classifying the subtype — adductor (ADSD) or abductor (ABSD) — determines which muscle is injected with botulinum toxin and is the single most important step before any treatment begins.
Spasmodic dysphonia is classified as a focal dystonia, part of a family of disorders (also including blepharospasm, cervical dystonia, and writer's cramp) caused by dysfunction of the basal ganglia–thalamocortical sensorimotor network rather than any structural laryngeal lesion. Functional imaging shows abnormal activation and impaired surround inhibition in the sensorimotor cortex, putamen, and cerebellum during phonation.
A genetic contributor is identified in a minority of cases — mutations in THAP1 have been linked to laryngeal dystonia, and roughly 12% of patients report an affected family member. In many patients, however, symptoms begin after an upper respiratory infection, a period of vocal strain, or emotional stress, suggesting the network dysfunction is unmasked by a peripheral trigger rather than caused by it.
At the muscle level, surface and needle EMG during spasms shows bursts of excessive, poorly modulated motor unit activity in the affected muscle group precisely timed to voicing attempts — this is what botulinum toxin ultimately targets.
Because the defect lies in central motor control rather than the muscle itself, botulinum toxin does not cure spasmodic dysphonia — it treats the peripheral output (excessive muscle contraction) and must be repeated for life.
Adductor SD (ADSD, ~85% of cases): involuntary hyperadduction of the vocal folds during voicing produces a strained, strangled, effortful voice quality with abrupt voice breaks on vowels, particularly in words beginning with vowels or voiced consonants. Speech is disproportionately affected compared with singing, laughing, shouting, or whispering — a hallmark task-specificity.
Abductor SD (ABSD, ~15% of cases): involuntary hyperabduction of the vocal folds during voiceless-to-voiced consonant transitions causes sudden breathy, aphonic breaks, typically on voiceless consonants (s, h, p, t, k) followed by voiced vowels. ABSD is generally harder to treat and to hear reliably on casual listening, contributing to diagnostic delay.
Mixed SD, combining both patterns, is rare and often the most difficult to manage.
Diagnosis is clinical, made by an experienced laryngologist and speech-language pathologist team: perceptual voice analysis during connected reading, spontaneous speech, and task-specific probes (singing, whispering) reveals the characteristic pattern-specific breakdown. Flexible fiberoptic laryngoscopy performed during connected speech directly visualizes the involuntary adduction or abduction spasms of the vocal folds in real time.
Acoustic analysis (voice breaks per minute, jitter, shimmer) can quantify severity and track treatment response. Mimicking conditions must be excluded: muscle tension dysphonia, vocal fold paresis/paralysis, essential vocal tremor, and multiple sclerosis can all resemble SD and require different management.
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| Adductor SD (ADSD) | Strained, strangled, effortful voice; abrupt breaks on vowels | Thyroarytenoid (TA) muscle — bilateral, 1.25–3.75 U/side | ~85% of cases; injection via cricothyroid membrane; most predictable response |
| Abductor SD (ABSD) | Breathy, aphonic breaks on voiceless consonants | Posterior cricoarytenoid (PCA) muscle — usually unilateral, 2.5–15 U | ~15% of cases; technically harder EMG target; risk of airway compromise if bilateral |
The larynx is controlled by five paired intrinsic muscles plus the unpaired interarytenoid, each acting on the arytenoid or thyroid cartilages to open, close, or tense the vocal folds. Only two of these muscles are relevant to spasmodic dysphonia treatment: the thyroarytenoid and the posterior cricoarytenoid — and choosing the wrong one, or missing it, renders an injection useless or harmful.
Thyroarytenoid (TA): the bulk of the vocal fold itself, running from the thyroid cartilage anteriorly to the arytenoid vocal process posteriorly. Contraction shortens and bulges the fold, adducting and tensing it — the prime mover of the strained hyperadduction seen in ADSD.
Lateral cricoarytenoid (LCA) and interarytenoid (IA): accessory adductors that rotate the arytenoids medially and appose them, respectively, to seal the posterior glottis.
Posterior cricoarytenoid (PCA): originates on the posterior cricoid lamina and inserts on the arytenoid muscular process; it is the only muscle in the body that abducts (opens) the vocal folds, essential for breathing. Its involuntary hyperactivity during voicing produces the breathy breaks of ABSD.
Cricothyroid (CT): tilts the thyroid cartilage forward relative to the cricoid, lengthening and tensing the folds for pitch control — not a direct SD target but relevant to needle trajectory planning.
Because the PCA is the only glottis-opening muscle, weakening it in ABSD carries a categorically different risk than weakening the TA in ADSD: excessive bilateral PCA denervation can leave the airway unable to open, risking stridor or airway obstruction — which is why PCA injections are usually staged and unilateral.
The TA occupies the entire body of the vocal fold, making it both the source of pathological hyperadduction and a relatively accessible, reliable injection target. It is reached percutaneously through the cricothyroid membrane — a thin, avascular window just below the thyroid cartilage — with the needle angled superiorly and laterally roughly 20–30° off the midline to enter the muscle belly on the target side, sometimes bilaterally.
The PCA sits on the posterior surface of the cricoid lamina, largely hidden behind the thyroid ala and adjacent to the pharyngoesophageal wall and the recurrent laryngeal nerve. It cannot be reached through the same anterior cricothyroid window; instead a posterolateral or transcartilaginous approach is used, typically guided by EMG feedback during a sniff maneuver (which selectively activates the PCA, unlike phonation which activates the adductors). Because of the proximity to the airway lumen and the disproportionate consequence of weakening the only abductor, PCA dosing is more conservative and individualized than TA dosing.
Precision matters enormously in laryngeal botulinum toxin injection: the target muscles are only millimeters from the airway lumen, major vessels, and each other. Electromyography (EMG) guidance — listening to the electrical signature of muscle activity through the needle tip itself — is what makes accurate, reproducible dosing possible in a structure this small and mobile.
With the patient supine and neck slightly extended, the hollow EMG needle-electrode is inserted through the skin over the cricothyroid membrane at or near the midline and advanced posterosuperiorly and laterally toward the target vocal fold, typically 3–5 mm off midline and a similar depth. The patient is asked to phonate a sustained vowel ("eee"). Correct placement within the TA produces crisp, high-amplitude, rapidly firing motor unit action potentials time-locked to phonation; a flat or distant signal means the needle must be redirected before any toxin is given.
For the PCA, the needle is passed lateral and posterior to the cricoid cartilage, often through a transcartilaginous route through the cricoid lamina itself, aiming for the posterior cricoid surface. Because the PCA activates on inspiration rather than phonation, the patient is asked to sniff sharply; a burst of motor unit activity synchronized with the sniff — and silence during phonation — confirms the needle tip is in the PCA rather than an adjacent adductor muscle.
EMG signal quality is binary in practice: crisp, phase-locked motor unit potentials mean "inject here"; polyphasic, distant, or absent signals mean "redirect the needle" — toxin is never given based on anatomic landmarks alone.
Typical starting doses are deliberately conservative: 1.25–3.75 U of onabotulinumtoxinA per vocal fold for ADSD (bilateral, or occasionally a single higher unilateral dose of up to ~5 U), versus 2.5–15 U for ABSD (typically unilateral, staged over separate visits if bilateral weakening is eventually needed). Dose is individualized over successive treatments based on the patient's prior response, duration of benefit, and side-effect tolerance — most laryngologists adjust up or down by increments of 20–25% between visits until an optimal steady state is reached.
Botulinum toxin type A is one of the most potent biological toxins known, yet its precision at the molecular level is what makes it a safe, reversible therapeutic tool. A single, well-characterized enzymatic cleavage event inside the nerve terminal is responsible for the entire clinical effect.
The BoNT/A holotoxin is a 150 kDa dichain protein: a 100 kDa heavy chain and a 50 kDa light chain joined by a single disulfide bond. The heavy chain's C-terminal domain binds with high specificity to a dual receptor complex on the presynaptic cholinergic nerve terminal — the synaptic vesicle protein SV2 together with a polysialoganglioside (GT1b) embedded in the membrane. This dual requirement gives the toxin its exquisite selectivity for motor and autonomic nerve terminals. Binding triggers receptor-mediated endocytosis, pulling the entire toxin into an intracellular vesicle.
As the endocytic vesicle acidifies, the heavy chain N-terminal domain inserts into the vesicle membrane and forms a channel, through which the light chain is translocated (unfolded) into the cytosol and refolds into its active conformation. The light chain is a zinc-dependent endopeptidase that recognizes and cleaves a single, specific peptide bond in SNAP-25 (synaptosomal-associated protein of 25 kDa) — one of the three core SNARE proteins, alongside syntaxin-1 and VAMP/synaptobrevin, that normally zipper together to force a synaptic vesicle membrane to fuse with the nerve terminal membrane.
Truncated SNAP-25 can no longer support productive SNARE-complex assembly.
Without an intact SNARE complex, acetylcholine-containing vesicles docked at the active zone cannot fuse with the terminal membrane — even though the nerve impulse still arrives normally, no neurotransmitter is released, and the postsynaptic muscle end-plate potential never reaches threshold.
The result is a graded, dose-dependent chemical denervation: the affected muscle fibers receive nerve impulses but cannot contract normally because acetylcholine release is blocked at the neuromuscular junction. This is not a destructive or permanent injury — the nerve terminal, axon, and muscle fiber remain structurally intact. Over subsequent weeks, the terminal begins sprouting new, unblocked synaptic contacts, gradually restoring transmission; the original terminal itself typically recovers and the sprouts are pruned once function returns, which is why the clinical effect fades over months rather than persisting indefinitely.
Botulinum toxin does not cure spasmodic dysphonia, but for most patients it converts a severely disabling, socially isolating voice disorder into a manageable chronic condition — at the cost of a lifelong cycle of re-injection every few months to sustain benefit.
Immediately after injection there is no change — the toxin must first bind, internalize, and cleave SNAP-25 inside the nerve terminal, a process that takes 1–4 days before any weakness is clinically apparent. Weakness then builds over the first week as more neuromuscular junctions become blocked. Voice quality typically peaks around 2 weeks, sometimes preceded by several days of transient breathiness or a whispery quality as the dose finds its balance point between reduced spasm and reduced strength. Benefit then plateaus for several weeks to a couple of months before gradually waning as nerve terminal sprouting restores transmission, prompting the patient to return for re-injection before symptoms fully return.
Because the toxin acts locally but can diffuse a few millimeters beyond the target muscle, side effects are direct extensions of its mechanism rather than off-target toxicity: breathy, weak voice (hypophonia) occurs in a dose-dependent minority of ADSD patients, usually resolving within 1–3 weeks. Mild dysphagia, particularly to thin liquids, occurs in roughly 15–30% of TA injections due to toxin diffusing to nearby pharyngeal constrictors, and is typically transient. For ABSD, the principal concern is airway compromise from excessive bilateral PCA weakness, which is why bilateral PCA treatment is usually staged across separate sessions rather than injected simultaneously.
Because the underlying basal ganglia network dysfunction is not altered by peripheral chemodenervation, treatment is repeated indefinitely — typically every 3–4 months for decades. Over repeated treatments, laryngologists titrate dose, injection site, and unilateral-versus-bilateral strategy to maximize the duration of good voice quality while minimizing breathy or dysphagic side effects, and most patients maintain a stable, individualized dosing regimen that provides consistent benefit for many years.
Longitudinal studies of botulinum toxin therapy for spasmodic dysphonia — now used clinically since the 1980s — show sustained efficacy without loss of response (antibody-mediated resistance is rare at the low doses used in laryngeal injection), making it the established first-line treatment worldwide.