Lexicon Development for COVID-19-related Concepts Using Open-source Word Embedding Sources: An Intrinsic and Extrinsic Evaluation

Soham Parikh
Anahita Davoudi
Shun Yu
Carolina Giraldo
Emily Schriver
Danielle Mowery

This article has been Reviewed by the following groups

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

Evaluated articles (ScreenIT)

Abstract

Scientists are developing new computational methods and prediction models to better clinically understand COVID-19 prevalence, treatment efficacy, and patient outcomes. These efforts could be improved by leveraging documented COVID-19–related symptoms, findings, and disorders from clinical text sources in an electronic health record. Word embeddings can identify terms related to these clinical concepts from both the biomedical and nonbiomedical domains, and are being shared with the open-source community at large. However, it’s unclear how useful openly available word embeddings are for developing lexicons for COVID-19–related concepts.

Objective

Given an initial lexicon of COVID-19–related terms, this study aims to characterize the returned terms by similarity across various open-source word embeddings and determine common semantic and syntactic patterns between the COVID-19 queried terms and returned terms specific to the word embedding source.

Methods

We compared seven openly available word embedding sources. Using a series of COVID-19–related terms for associated symptoms, findings, and disorders, we conducted an interannotator agreement study to determine how accurately the most similar returned terms could be classified according to semantic types by three annotators. We conducted a qualitative study of COVID-19 queried terms and their returned terms to detect informative patterns for constructing lexicons. We demonstrated the utility of applying such learned synonyms to discharge summaries by reporting the proportion of patients identified by concept among three patient cohorts: pneumonia (n=6410), acute respiratory distress syndrome (n=8647), and COVID-19 (n=2397).

Results

We observed high pairwise interannotator agreement (Cohen kappa) for symptoms (0.86-0.99), findings (0.93-0.99), and disorders (0.93-0.99). Word embedding sources generated based on characters tend to return more synonyms (mean count of 7.2 synonyms) compared to token-based embedding sources (mean counts range from 2.0 to 3.4). Word embedding sources queried using a qualifier term (eg, dry cough or muscle pain) more often returned qualifiers of the similar semantic type (eg, “dry” returns consistency qualifiers like “wet” and “runny”) compared to a single term (eg, cough or pain) queries. A higher proportion of patients had documented fever (0.61-0.84), cough (0.41-0.55), shortness of breath (0.40-0.59), and hypoxia (0.51-0.56) retrieved than other clinical features. Terms for dry cough returned a higher proportion of patients with COVID-19 (0.07) than the pneumonia (0.05) and acute respiratory distress syndrome (0.03) populations.

Conclusions

Word embeddings are valuable technology for learning related terms, including synonyms. When leveraging openly available word embedding sources, choices made for the construction of the word embeddings can significantly influence the words learned.

SciScore for 10.1101/2020.12.29.20249005: (What is this?)

Please note, not all rigor criteria are appropriate for all manuscripts.

Table 1: Rigor

NIH rigor criteria are not applicable to paper type.

Table 2: Resources

Software and Algorithms
Sentences	Resources
We also depict semantic disagreements between each pair of annotators using heatmaps generated using matplotlib [34].	matplotlib suggested: (MatPlotLib, RRID:SCR_008624)

Results from OddPub: Thank you for sharing your code.

Results from LimitationRecognizer: We detected the following sentences addressing limitations in the study:

Limitations and Future Work: Our study has a few notable limitations. We began this study during the early stages of the COVID-19 pandemic when the symptomatology was less understood. COVID-19 is a heterogeneous disease with emerging symptomatology identified …

SciScore for 10.1101/2020.12.29.20249005: (What is this?)

Please note, not all rigor criteria are appropriate for all manuscripts.

Table 1: Rigor

NIH rigor criteria are not applicable to paper type.

Table 2: Resources

Software and Algorithms
Sentences	Resources
We also depict semantic disagreements between each pair of annotators using heatmaps generated using matplotlib [34].	matplotlib suggested: (MatPlotLib, RRID:SCR_008624)

Results from OddPub: Thank you for sharing your code.

Results from LimitationRecognizer: We detected the following sentences addressing limitations in the study:

Limitations and Future Work: Our study has a few notable limitations. We began this study during the early stages of the COVID-19 pandemic when the symptomatology was less understood. COVID-19 is a heterogeneous disease with emerging symptomatology identified through ongoing clinical observational studies. Emerging COVID-19-related symptomatology, i.e., loss of smell and loss of taste and COVID toes were not included in our analysis as their association with COVID-19 were not well-understood at the time of our study. We leveraged existing word embedding sources to better understand the utility of embeddings for synonym generation. We recognize that further experimentation is needed to support broader claims of their utility. As a proof-of-concept of patient information retrieval, we applied an expanded lexicon of terms representing clinical features of COVID-19 to three disorder cohorts (pneumonia, ARDS, and COVID-19). Although these terms retrieved a high proportion of patients, we acknowledge that additional terms might be necessary to accurately identify these features and that contextualization (i.e., negation, severity, experiencer, temporality [37]) is critical to generating accurate patient profiles. We look forward to addressing these issues as next steps toward developing our clinical information extraction pipeline in our follow-up study.

Results from TrialIdentifier: No clinical trial numbers were referenced.

Results from Barzooka: We did not find any issues relating to the usage of bar graphs.

Results from JetFighter: We did not find any issues relating to colormaps.

Results from rtransparent:

Thank you for including a conflict of interest statement. Authors are encouraged to include this statement when submitting to a journal.
Thank you for including a funding statement. Authors are encouraged to include this statement when submitting to a journal.
No protocol registration statement was detected.

Read the original source

Version published to 10.2196/21679
Feb 22, 2021
Version published to 10.1101/2020.12.29.20249005 on medRxiv
Jan 4, 2021
Version published to 10.2196/preprints.21679
Jun 21, 2020

A Pilot Project Leveraging Large Language Models for Automated Screening and Variable Extraction in Observational Studies

This article has 11 authors:
1. Manjil M. Pradhan
2. Rajesh Upadhayaya
3. Sarah C. Wenyon
4. Alexandria Viszolay
5. Melissa Rethlefsen
6. Vincent Metzger
7. Gerardo Villarreal
8. Santiago Alvarez Lesmes
9. David Andrade
10. Jason Timm
11. Scott A. Malec
This article has no evaluationsLatest version Jun 15, 2026
Relationship Extraction for Adverse Drug Events in Clinical Notes Using Large Language Models

This article has 10 authors:
1. Joseph M Plasek
2. Yiming Li
3. Mary G Amato
4. Dinah Foer
5. Diane L. Seger
6. Shayma Alzaidi
7. Huiyuan Zhou
8. Gretchen Purcell Jackson
9. David W Bates
10. Li Zhou
This article has no evaluationsLatest version Jun 1, 2026
PrimeKG-Plus: a refreshed and rare-disease-enriched precision medicine knowledge graph

This article has 12 authors:
1. Trinh Trung Duong Nguyen
2. Thuy Nguyen-Phuong
3. Quy-Hoai Nguyen
4. Amna Mumtaz Abbasi
5. Hanh-Dung Le Phan
6. Luong Bao-Anh Nguyen
7. Nhat-Thien Phan
8. Nurettin Nusret Curabaz
9. Alexander S. Hauser
10. Ziaurrehman Tanoli
11. Dinh Truong Nguyen
12. Albert J. Kooistra
This article has no evaluationsLatest version Jul 18, 2026

This article has been Reviewed by the following groups

Discuss this preprint

Listed in

Abstract

Objective

Methods

Results

Conclusions

Article activity feed

Related articles

A Pilot Project Leveraging Large Language Models for Automated Screening and Variable Extraction in Observational Studies

Relationship Extraction for Adverse Drug Events in Clinical Notes Using Large Language Models

PrimeKG-Plus: a refreshed and rare-disease-enriched precision medicine knowledge graph