The McGurk Effect: Why What You See Changes What You Hear
Close your eyes and you hear one syllable. Open them and you hear another.

Play the sound "ba" over a video of a mouth saying "ga" and many people hear "da" — a syllable present in neither channel. The McGurk effect shows that what you hear is assembled from sound and vision together. What almost every popular account leaves out is that susceptibility varies enormously: in one large study, some participants perceived the illusion on essentially every trial and others on none at all.
That variability is not a footnote. It is the reason a growing number of researchers argue the effect is being used to measure something it does not reliably measure.
What McGurk and MacDonald found
Harry McGurk and John MacDonald published the effect in Nature in 1976, and by their own account discovered it by accident while dubbing speech onto video for a study about something else.
The setup is simple. An auditory syllable is paired with a video of a face articulating a different one. The most-used pairing is auditory /ba/ with visual /ga/, which commonly produces the perception of /da/ — a compromise between what the ears receive and what the eyes see.
The striking part is that it is not a judgement you can override. Knowing exactly how the illusion works does not stop it. Close your eyes and you hear the true syllable; open them and the perception changes back.
Why vision gets a vote at all
Speech perception evolved face to face. Lip and jaw movements carry real information about which sounds are being produced, and that information is especially valuable when the acoustic signal is degraded — a noisy room, a poor connection, a quiet talker.
So the auditory system does not simply decode sound. It integrates the acoustic signal with visual articulatory information into a single best estimate, with the superior temporal sulcus repeatedly implicated as a site where the two streams meet.
When the two channels agree, this is invisible and useful. When an experimenter makes them disagree, the integration is exposed, and the output is a syllable neither channel supplied.
It is the same broad principle behind why your recorded voice sounds wrong: what you perceive as speech is a construction from multiple inputs, not a readout of one.
The part popular coverage omits
Demonstrations imply the effect is close to universal. It is not.
Debshila Basu Mallick, John Magnotti and Michael Beauchamp tested 165 English-speaking adults on twelve different McGurk stimuli. Susceptibility ranged across the entire scale, and the distribution was strongly bimodal: around 77 percent of participants either almost never perceived the illusion or almost always did.
The rate also depends heavily on which stimulus is used — reported fusion rates vary from single figures to the high seventies across different recordings of the same syllable pairing. One later study across 199 participants found an overall fusion rate of about 20 percent, far below the impression given by the original demonstrations.
It is stable, which is the interesting part
Whatever determines your susceptibility, it is a property of you rather than of the moment. Forty participants retested a year later showed high test-retest reliability — the people who saw the illusion still saw it, and those who did not still did not.
That makes it a genuine individual difference, like frisson or the photic sneeze reflex, rather than measurement noise. What produces it is not settled; the best-supported contributor so far is fine-grained lipreading ability rather than any general attentional or working-memory measure.
Why researchers are backing away from it
The McGurk effect became a standard assay for audiovisual speech integration — used to compare groups, populations and clinical conditions. Several researchers now argue that this was a mistake.
The objection is specific. If susceptibility ranges from zero to a hundred percent among typical adults, and does not correlate well with other measures of audiovisual speech integration, then a group difference in McGurk rates may reflect sampling rather than integration. Magnotti and Beauchamp pointed out that reported differences between autistic and non-autistic groups have run from 45 percent lower to 10 percent higher across studies — a spread that suggests the measure, not the groups.
None of this means the effect is not real. It is one of the most reliably reproducible illusions in perception at the group level. The problem is using an individual's score as a number that means something.
What makes it stronger or weaker
If a demonstration does nothing for you, the conditions matter more than most people realise. Degraded audio increases the effect, because the brain leans harder on vision when the sound is unreliable. Clean, loud audio in a quiet room reduces it.
So does believing the face and voice belong together. Work using a male face with a female voice found that familiarising participants with the true pairings reduced susceptibility — once the brain has reason to treat the two streams as separate sources, it integrates them less.
Which is a useful reframe. The illusion is not a glitch to be resisted; it is the system correctly weighting an unreliable channel against a more reliable one, using an assumption that these two signals come from the same mouth.
What to take from it
That hearing is not a purely auditory sense is the durable finding, and it holds regardless of how susceptible you personally are. Vision shapes what you hear during every conversation you have; the illusion just makes it visible by forcing a disagreement.
The practical version is unremarkable and real: you understand speech better when you can see the speaker, which is why a noisy restaurant is harder on a phone call than in person, and why watching someone's face genuinely helps when the audio is poor.
It also fits a pattern this site keeps running into — speech is not a channel but a whole-body act, which is why people gesture on the phone to listeners who cannot see them.
Same trick, different system
The Dive Reflex: Why Cold Water on Your Face Slows Your Heart
Frequently asked questions
An illusion in which an auditory syllable paired with a video of a mouth saying a different one produces the perception of a third syllable. Auditory /ba/ with visual /ga/ commonly yields /da/ — a compromise between what the ears and eyes report.
Susceptibility varies enormously and is a stable individual trait. In a study of 165 adults, around 77 percent either almost never or almost always perceived the illusion, and rates ranged across the full scale from zero to a hundred percent.
No. Understanding the mechanism does not prevent it. Perception updates automatically when the visual channel is available — closing your eyes restores the true syllable, and opening them changes it back.
Increasingly disputed. Because individual susceptibility ranges from zero to a hundred percent and correlates poorly with other integration measures, group differences may reflect sampling rather than the ability being measured.
Lip and jaw movements carry genuine information about which sounds are being produced, and that information is most useful when the acoustic signal is degraded. The brain integrates both streams into one estimate, which is normally invisible because they usually agree.
This article is educational science trivia about everyday human biology and psychology. It is not medical advice, diagnosis, or treatment, and it is not a substitute for care from a qualified professional.


