Hearing with your eyes
Play the audio "ba" while showing a video of a mouth clearly articulating "ga", and most listeners report hearing a third sound entirely — "da" — that matches neither the audio nor the video track alone. Discovered by Harry McGurk and John MacDonald in 1976 (originally by accident, while dubbing mismatched infant-directed speech for an unrelated study), the McGurk effect is one of psychology's most robust demonstrations that speech perception is inherently multisensory: the brain doesn't just hear phonemes, it integrates sound with visible articulatory movement, and it does this automatically, even in people who know exactly what's happening and are trying not to be fooled.
Why the fusion lands on "da"
The specific fused percept is not arbitrary. Consonants are partly defined by place of articulation — where in the mouth the closure happens. /b/ is bilabial (lips together), /g/ is velar (back of tongue against the soft palate), and /d/ is alveolar (tongue against the ridge behind the upper teeth) — acoustically and visually intermediate between the other two. When the ears report a bilabial cue and the eyes report a velar cue, the brain's most likely inference given both streams together often lands on the intermediate alveolar place, producing the classic /b/-audio + /g/-video → /d/-perceived fusion. Other audiovisual pairings produce combination responses instead of fusion (e.g. audio /g/ + video /b/ is often heard as "bga"), which is one reason researchers treat the McGurk effect as evidence for genuine perceptual integration rather than the brain simply picking whichever channel it trusts more.
Bayesian cue combination
The modern computational framing treats speech perception as Bayesian cue combination: the brain holds noisy, uncertain estimates of the spoken phoneme from the auditory channel and from the visual channel separately, and combines them weighted roughly by each channel's reliability — a formalisation close to the general principle of optimal multisensory integration demonstrated across many perceptual domains (visual-haptic size and location judgements included), where more reliable senses are weighted more heavily and the combined percept has lower uncertainty than either sense alone.
fused percept ∝ likelihood(audio | phoneme) × likelihood(video | phoneme)
× prior(phoneme)
when audio strongly implies /b/ and video strongly implies /g/,
neither channel dominates outright — the posterior often peaks at the
intermediate place of articulation, /d/, rather than at either input
It survives knowing the trick
Unlike many cognitive illusions that weaken once explained, the McGurk effect persists almost undiminished even when a listener is told exactly what's happening and is actively trying to hear the true audio syllable — strong evidence that audiovisual speech fusion happens at an early, largely automatic stage of perceptual processing rather than at a late, cognitively-penetrable stage of deliberate judgement. Closing your eyes, however, removes it instantly, since it depends entirely on visual articulatory input being present and attended to.
Not universal, and clinically informative
Susceptibility to the McGurk effect varies by language, culture and age (it's typically weaker in young children and in some tonal-language populations, and differs across languages with different phoneme inventories and different amounts of native audiovisual speech exposure), and its strength has become a research tool in its own right: reduced McGurk susceptibility is studied as a marker of altered audiovisual integration in autism spectrum conditions, and the effect is used more broadly to probe how and when the developing brain starts binding auditory and visual speech cues into one percept, since infants show reduced fusion compared with adults.
Frequently asked questions
What exactly is the McGurk effect an illusion of?
It is an illusion of speech perception, not of hearing or vision in isolation. Mismatched auditory and visual syllables (e.g. audio "ba" dubbed onto a video mouth clearly saying "ga") fuse into a third percept, commonly "da," demonstrating that the brain integrates sound and lip movement automatically rather than treating them as separate channels.
Can you train yourself to stop experiencing the McGurk effect?
Not reliably. The fusion persists even in listeners who know exactly which syllable was spoken in the audio and are actively trying to report it correctly, which is taken as evidence that audiovisual speech integration happens early and automatically in perceptual processing rather than as a late, correctable judgement.
Why does mismatched ba/ga usually get perceived as da specifically?
Consonants differ by place of articulation: /b/ is produced with the lips (bilabial), /g/ with the back of the tongue (velar), and /d/ with the tongue against the ridge behind the teeth (alveolar) — acoustically and visually between the other two. When the ears suggest bilabial and the eyes suggest velar, the brain's best combined estimate often lands on the intermediate alveolar place, producing "da."
Try it live
Everything above runs in your browser — open McGurk Effect and change the parameters while it is running. Nothing is installed, nothing is uploaded, the whole model lives in one tab.
▶ Open McGurk Effect simulation