Articulatory timing and form support distinct neural benefits during audiovisual speech

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

In noisy environments, visible speech articulations improve listening comprehension. The benefit derives from several sources, including articulatory timing and shape. Recent research has shown that visual cortex encodes a categorical representation of articulatory features and that visual speech can benefit both acoustic and phonetic feature processing separately. The present study advances the hypothesis that the shape of the articulators specifically influences the categorization of auditory speech in terms of its phonetic features. We tested this by linearly modeling electroencephalographic responses to natural, continuous speech (in noise) in terms of the acoustic and articulatory features of the speech. We compared the performance of these models in conditions where the speech was accompanied by a natural video of the speaker with their mouth visible, and a video where their mouth was covered by a dynamic ellipse obscuring articulatory shape but preserving dynamics. The dynamic mask reduced comprehension, neural processing of phonetic features, the associated multisensory benefits, and indices of visual-only linguistic processing over occipital scalp. Our findings support substantial visual involvement in speech comprehension, derived largely from the shape of the articulators. They also corroborate several proposals involving audiovisual speech processing hierarchy and the nature of the information contained in visible speech.

Highlights

  • Visual speech provides at least two forms of information to enhance acoustic speech processing: redundant temporal dynamics and complementary articulatory information

  • Covering the mouth with a dynamic mask preserves horizontal and vertical lip movement information, but largely removes articulatory detail

  • Visual speech with a mask preserves some general multisensory benefits but removes visual linguistic information and its ability to enhance auditory processing at the level of phonetic features.

Article activity feed