Semi-automated annotation refinement accelerates cell type identification in brain spatial and single-cell studies
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background
single-cell and spatial omic techniques have enabled the investigation of cell type specific alterations in biologically complex tissues. In an effort to map cell taxonomies, large atlas-based studies and multi-laboratory consortia have created sets of annotated cell types. However, application of atlas- or database-level knowledge to individual studies is often resource-limited and computational demands scale with the size of both query and reference datasets.
Results
Here, we report a statistical framework for rapid label transfer using summary statistics and user-defined hyperparameters. Semi-Automated Hand Annotation (SAHA) 1 allows the user to investigate magnitude, directionality, and statistical significance of matches between unnamed query clusters and reference cell types using either marker-based or marker-free methodologies. By pre-loading the package with summary statistics from the Allen Brain Cell Atlas of the mouse brain, the SAHA R package is capable of rapid cell type comparisons that closely mimic cell typing by integration-based annotation strategies. Furthermore, this flexible package is capable of comparisons across omic modalities, cluster resolutions, and annotations from any study where summary statistics are available. We demonstrate this flexibility by using multiple single-nuclei studies of the mouse cerebellum, mouse cerebral cortex, human cerebral cortex, human peripheral blood mononuclear cells, and one mouse spatial transcriptomic assay. Importantly, this method avoids privacy concerns as it does not require the sharing or deposition of raw data in a web-based tool.
Conclusions
As a result, SAHA offers a non-deterministic annotation reporting structure with automated html reports and summary statistics for transparency in cell typing decisions. Taken together, this scalable framework implemented as a package in R affords increased biological insight into the annotation of single-cell and spatial datasets.
SHORT SUMMARY
Acri and colleagues present rapid cell type annotation without the need for dataset integration. This paper outlines the utility of the package, SAHA, in annotating neurological datasets.