Large-Scale Psychiatric Concept Extraction from Electronic Health Records: A Comparative Study of Encoder-Based Language Models

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Free-text notes in electronic health records (EHRs) contain fine-grained psychiatric information that is essential for psychiatric research and clinical care, and often absent or under-recorded in structured codes alone. Clinical natural language processing (cNLP) can support extraction of this information from EHR notes, yet Spanish-language cNLP remains under-developed. Moreover, broad evaluations comparing multiple encoder-based language models across extensive, fine-grained psychiatric concept sets remain scarce, and it remains unclear how these models compare with traditional NLP (tNLP) systems and much larger generative large language models (LLMs). In addition, cross-site performance of fine-tuned models is rarely tested, and limited annotated training data remains a major challenge, especially for rare symptoms.

Objectives

We aimed to advance scalable, global psychiatric cNLP by fine-tuning multiple encoder-based models with differing architectures and pre-training strategies for detecting fine-grained psychiatric concepts in Spanish EHRs. We further evaluated the impact of augmenting the fine-tuning data with precision-weighted weak labels for less-frequent concepts, and compared the performance of the encoder-based models to that of tNLP and a fine-tuned generative LLM trained on the same data. Finally, we evaluated model cross-site generalizability on an external EHR dataset.

Methods

Three encoder-based models (BETO, XLM-RoBERTa-large, and bsc-bio-ehr-es) were fine-tuned on 1,642 clinician-annotated EHR documents from Colombia to detect 110 psychiatric concepts in Spanish text. To address the limited annotated examples available for less-frequent concepts, 12,000 additional documents were weakly-labeled for less-frequent concepts using tNLP, and incorporated into the fine-tuning data with labels weighted by pattern precision. Models were compared with tNLP and a generative LLM, and evaluated on an external EHR dataset from another psychiatric hospital in Colombia.

Results

Encoder model performance varied substantially, with macro-F1 ranging from 0.64 to 0.81. BETO achieved the highest macro-F1 (0.81; median F1=0.88 [IQR=0.77-0.96]). Adding precision-weighted weak labels for less-frequent concepts improved BETO’s overall macro-F1 to 0.83 and increased mean F1 for the 55 augmented concepts from 0.82 to 0.86. Under matched fine-tuning conditions, fine-tuned BETO and the tNLP method were equivalent in F1, whereas the LLM significantly outperformed BETO in F1. After weak-label augmentation, BETO significantly outperformed tNLP in F1 ( P FDR <.001) and narrowed the performance gap with the LLM, although equivalence was not established. Lastly, fine-tuned BETO maintained reasonably strong performance on data from an external hospital not used for model fine-tuning (out-of-domain macro-F1=0.78).

Conclusions

General-purpose pre-trained encoders had strong performance for psychiatric concept extraction from Spanish EHRs. Weak-label augmentation improved BETO’s performance and strengthened results relative to a tNLP baseline, while reducing, but not eliminating, the performance gap with a much larger fine-tuned generative LLM. These findings highlight the utility of these relatively lightweight models for scalable, accurate and reproducible detection of psychiatric concepts in Spanish-language EHRs.

Article activity feed