Visual and auditory deep learning models capture neural representations of naturalistic social interaction in the superior temporal sulcus

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

The superior temporal sulcus (STS) is selectively responsive to multimodal social interactions. Yet studies so far have relied on pre-defined, simplified stimuli or features to uncover the type of information that drives STS activity. We hypothesized that high-dimensional representations from visual and auditory deep learning models would better predict STS responses to naturalistic social interactions. We used self-supervised visual and auditory deep learning models to extract representations of movie frames and audio, respectively, of a TV series participants watched during fMRI scanning. Voxel-wise encoding models of a joint visual-auditory representation outperformed human-made social-affective annotations in predicting STS. Variance partition further revealed visual-auditory posterior-to-anterior gradient within the STS. To interpret what these models encode, we applied Principal Component Analysis to the encoding model weights. In both the visual and auditory models the first dimension tracked social interaction and peaked in the STS, indicating that social interaction is a dominant dimension of STS representation across both modalities. We conclude that the STS represents naturalistic social interaction in a multimodal manner, integrating visual and auditory information, and that visual and auditory deep learning models capture key representational properties of these responses.

Article activity feed