Task-dependent model selection for structured extraction from multilingual non-English clinical records
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Objective
To evaluate task-dependent selection of local and cloud models for structured extraction from multilingual non-English stroke discharge summaries, distinguishing entity detection, record assembly, and raw-document processing.
Methods
This retrospective system evaluation compared multilingual encoders, locally fine-tuned Qwen3-4B models, and zero-shot GPT-5.5. Primary test cohorts comprised 332 section cases, 149 medication cases with 2,475 reference records after identity-based exclusions, and 191 laboratory cases. Outcomes were section-span F1, drug-name F1, normalized seven-field record recovery, test-name F1, and laboratory quintuple F1. Paired human comparisons used a common second-annotator reference on 50 cases per task. Saved cascade outputs were compared with curated-section controls using document-bootstrap intervals.
Results
Section F1 was 0.919 for the encoder, 0.926 for GPT-5.5, and 0.932 for Qwen. GPT-5.5 led medication detection (0.966 versus 0.940 for Qwen), whereas Qwen led normalized medication recovery (0.381 versus 0.311; difference 0.070, 95% CI 0.026–0.114) and laboratory quintuple F1 (0.892 versus 0.822). With matched inference stacks, medication recovery declined from 0.377 on curated inputs to 0.204 for single-window and 0.246 for all-block cascades; quintuple F1 on the original laboratory cohort declined from 0.898 to 0.811. Corpus-wide extraction processed 193,101 summaries in 13.3 H100 GPU-hours.
Conclusion
Local models achieved strong detection performance without frontier-scale inference in this setting. Fine-tuned local LLMs were advantageous for record assembly under the evaluated recipes; annotation conventions, normalization, and upstream section extraction remained important constraints.