Inferring Protein Variant Impacts Across Contexts

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Multiplexed assays of variant effects (MAVEs) measure the functional impact of many protein sequence variants in parallel, potentially covering all possible single amino acid substitutions. Unlike current computational variant effect predictors, MAVEs can reveal the effects of variants under different genetic and environmental contexts. However, whereas the space of possible contexts is effectively infinite, ‘contextual’ MAVE studies are limited by finite experimental budgets. To maximize coverage across contexts, one strategy is to carry out sub-saturation contextual MAVEs and then fill in the gaps via imputation. Here we categorize and compare different imputation challenges, explore a collection of multi-context imputation solutions (including linear mixed-effects models, random forests, and autoencoders), and provide insight into how best to proceed in any given imputation task. We find that which method is optimal depends on the imputation task and on how densely the contexts have been measured, with more flexible models excelling when measurements are plentiful and the simplest models proving most reliable when measurements are sparse. However, the simplest models, though well suited to imputing scores where variants have been measured in the source context, cannot impute scores where variants were not measured in either context. This proves to be a major limitation when both maps are sparsely measured. We provide a conceptual framework and an initial evaluation of multi-context imputation methods that can extend the scope of large-scale studies of context-dependent variant effects.

Article activity feed