Characterization of Influenza A HA and NA Subtypes Using Protein Sequences

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Motivation

Influenza is an infectious disease associated with excess human deaths. For influenza A virus (IAV), the surface antigens hemagglutinin (HA) and neuraminidase (NA) define the specific subtype. The high mutation rate of IAV can make clinical testing and assigning subtype after sequencing challenging. Accurate subtype classification is important for tracking circulating IAV strains that may impact human health.

Results

We analyzed a large IAV protein sequence dataset with known HA and NA subtypes. Using logistic regression and random forest approaches, we identified a small set of subtype-associated amino acids, which were then used to develop a subtype characterization method with near-optimal accuracy. Further analysis indicated that signal sequence peptide variation and indels in HA and NA explained the unique combination of subtype-associated amino acids.

Article activity feed