Adverse Drug Events Across Data-Production Contexts: Multilingual Detection, Alignment, and Cross-Genre Discourse Analysis
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Adverse drug event (ADE) evidence is produced across patient-generated, clinical, and scientific settings that differ in language, documentation purpose, terminology, and degree of standardization. These differences shape both which adverse experiences become visible to pharmacovigilance systems and how readily they can be linked to curated drug-safety knowledge. We examine these relationships across five corpora representing distinct data-production settings: ADE Corpus V2 (medical case reports), SMM4H-2026 Task 1 (multi-lingual user-generated health content), CADEC V2 (patient-forum narratives), the Dutch ADE Corpus (EHR clinical notes), and TwiMed-PubMed (biomedical literature).
A shared BERTopic analysis of ADE-positive texts concerning antidepressants and antihypertensives across the four English-language corpora identified nine interpretable topics. CADEC V2 contained a more differentiated distribution of symptom-specific themes, including sexual effects, suicidal or panic-related thoughts, vivid dreams, and memory difficulties, whereas SMM4H-2026, TwiMed-PubMed, and ADE Corpus V2 were dominated by a broader medication, sleep, tiredness, and pain theme. These patterns indicate that data-production context shapes what adverse experiences are expressed and standardized, with patient-generated narratives surfacing subjective, symptom-specific experience largely absent from clinical and scientific sources.
We further show that this context shapes how readily real-world drug mentions can be linked to curated pharmacovigilance knowledge. Using SIDER 4.1 as a retrieval resource, we find substantial cross-corpus mismatches between real-world drug mentions and SIDER’s predominantly English, generic-name vocabulary: CADEC V2 achieved only 9.5% exact-match coverage, with unmatched mentions frequently involving brand names, misspellings, and language-specific variants, compared to 91.0% coverage in TwiMed-PubMed’s formally standardized biomedical literature.
To probe how these representational differences interact with automated detection, we compare corpus-specific QLoRA fine-tuning of Llama-3.2-3B with retrieval-augmented inference using Llama-3.1-70B and Llama-3.1-405B grounded in SIDER-retrieved evidence. QLoRA-Llama-3B achieved the highest micro-averaged F1 scores on ADE Corpus V2 (0.91), CADEC V2 (0.88), and SMM4H-2026 (0.80), whereas SIDER-grounded inference with Llama-3.1-405B achieved the highest scores on Dutch ADE (0.95) and TwiMed-PubMed (0.91); these corpus-dependent patterns should not be interpreted as a controlled comparison of adaptation strategies, since model scale, task formulation, and available supervision differ across datasets. Together, our findings indicate that data-production context influences what adverse experiences are expressed, how they are standardized, and how readily they can be retrieved and computationally detected. Pharmacovigilance systems should therefore combine source-sensitive supervision with external knowledge grounding while explicitly monitoring gaps between real-world language and curated drug-safety resources.