Characterization of long-term patient-reported symptoms of COVID-19: an analysis of social media data

Juan M. Banda

Nicola Adderley

Waheed-Ul-Rahman Ahmed

Heba AlGhoul

Osaid Alser

Muath Alser

Carlos Areia

Mikail Cogenur

Krisitina Fišter

Saurabh Gombar

Vojtech Huser

Jitendra Jonnagaddala

Lana YH Lai

Angela Leis

Lourdes Mateu

Miguel Angel Mayer

Evan Minty

Daniel Morales

Karthik Natarajan

Roger Paredes

Vyjeyanthi S. Periyakoil

Albert Prats-Uribe

Elsie G. Ross

Gurdas Singh

Vignesh Subbian

Arani Vivekanantham

Daniel Prieto-Alhambra

Abstract

As the SARS-CoV-2 virus (COVID-19) continues to affect people across the globe, there is limited understanding of the long term implications for infected patients ^1–3 . While some of these patients have documented follow-ups on clinical records, or participate in longitudinal surveys, these datasets are usually designed by clinicians, and not granular enough to understand the natural history or patient experiences of ‘long COVID’. In order to get a complete picture, there is a need to use patient generated data to track the long-term impact of COVID-19 on recovered patients in real time. There is a growing need to meticulously characterize these patients’ experiences, from infection to months post-infection, and with highly granular patient generated data rather than clinician narratives. In this work, we present a longitudinal characterization of post-COVID-19 symptoms using social media data from Twitter. Using a combination of machine learning, natural language processing techniques, and clinician reviews, we mined 296,154 tweets to characterize the post-acute infection course of the disease, creating detailed timelines of symptoms and conditions, and analyzing their symptomatology during a period of over 150 days.

Article activity feed

SciScore for 10.1101/2021.07.13.21260449: (What is this?)

Please note, not all rigor criteria are appropriate for all manuscripts.

Table 1: Rigor

NIH rigor criteria are not applicable to paper type.

Table 2: Resources

Software and Algorithms
Sentences	Resources
While the Twitter API only allows free users to extract a very limited set of old tweets, we used two scraping Python packages, GetOldTweets 34 and our own scraper part of SMMT35 to extract the user timelines.	Python suggested: (IPython, RRID:SCR_001658)
In order to normalize the annotations made by our clinicians, we used the Observational Health Data Sciences and Informatics (OHDSI) vocabulary that includes a collection of biomedical controlled vocabularies such as: RxNorm, SNOMED-CT, ICD9/10, CPT, among many others.	RxNorm suggested: (RxNorm, RRID:SCR_006645)

Results from OddPub: We did …

SciScore for 10.1101/2021.07.13.21260449: (What is this?)

Please note, not all rigor criteria are appropriate for all manuscripts.

Table 1: Rigor

NIH rigor criteria are not applicable to paper type.

Table 2: Resources

Software and Algorithms
Sentences	Resources
While the Twitter API only allows free users to extract a very limited set of old tweets, we used two scraping Python packages, GetOldTweets 34 and our own scraper part of SMMT35 to extract the user timelines.	Python suggested: (IPython, RRID:SCR_001658)
In order to normalize the annotations made by our clinicians, we used the Observational Health Data Sciences and Informatics (OHDSI) vocabulary that includes a collection of biomedical controlled vocabularies such as: RxNorm, SNOMED-CT, ICD9/10, CPT, among many others.	RxNorm suggested: (RxNorm, RRID:SCR_006645)

Results from OddPub: We did not detect open data. We also did not detect open code. Researchers are encouraged to share open data when possible (see Nature blog).

Results from LimitationRecognizer: An explicit section about the limitations of the techniques employed in this study was not found. We encourage authors to address study limitations.

Results from TrialIdentifier: No clinical trial numbers were referenced.

Results from Barzooka: We did not find any issues relating to the usage of bar graphs.

Results from JetFighter: We did not find any issues relating to colormaps.

Results from rtransparent:

Thank you for including a conflict of interest statement. Authors are encouraged to include this statement when submitting to a journal.
Thank you for including a funding statement. Authors are encouraged to include this statement when submitting to a journal.
Thank you for including a protocol registration statement.

Results from scite Reference Check: We found no unreliable references.

Read the original source

Laurence A. Jacobs

Clinical Study Protocol of the ‘Biomarkers of Severity of COVID-19 Patients’ (BIOMARCOVID) Project

Tuan-Anh Dinh
Corentin Leroy
Marion Brandolini-Bunlon
Sylvie Berthier
Candice Trocme
Nelle Varoquaux
Caroline Plazy
Antoine Vilotitch
Charles Terra
Bertrand Toussaint
Jean-Luc Bosson
Florence Castelli
Estelle Pujos-Guillot
Pauline Le Faouder
Justine Bertrand-Michel
Marion Le Marechal
Olivier Epaulard
Audrey Le Gouellec

A risk-of-contagion index using a Bayesian based model for the COVID-19 epidemic in Mexico

Ruth Corona-Moreno
M. Adrian Acuña-Zegarra
Mario Santana-Cibrian
Jorge X. Velasco-Hernandez

Characterization of long-term patient-reported symptoms of COVID-19: an analysis of social media data

This article has been Reviewed by the following groups

Discuss this preprint

Listed in

Abstract

Article activity feed

Pre-pandemic blood profiles predict COVID-19 hospitalization and death a decade later

Clinical Study Protocol of the ‘Biomarkers of Severity of COVID-19 Patients’ (BIOMARCOVID) Project

A risk-of-contagion index using a Bayesian based model for the COVID-19 epidemic in Mexico

This article has been Reviewed by the following groups

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Pre-pandemic blood profiles predict COVID-19 hospitalization and death a decade later

Clinical Study Protocol of the ‘Biomarkers of Severity of COVID-19 Patients’ (BIOMARCOVID) Project

A risk-of-contagion index using a Bayesian based model for the COVID-19 epidemic in Mexico