A framework for human-artificial intelligence co-learning for disease activity labeling using electronic health records

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Objective

To develop and evaluate a framework for human-AI interaction. This approach, SHARE (Synergistic Human-Agent REasoning system) was designed to support scalable phenotyping of complex outcomes accurately, robustly and reproducibly from real-world electronic health record (EHR) data to support real-world evidence (RWE) generation.

Methods and Analysis

Using rheumatoid arthritis (RA) disease activity as the use- case, we studied a multi-institutional EHR-based RA cohort of 3,167 patients. Expert reviewers and a disease activity agent labeled notes using the same review guideline. The agent combined embedding-based informative-note filtering, structured evidence extraction, and evidence-based integrated reasoning to assign disease activity categories with supporting evidence, rationale, confidence, and ambiguity flags. To support scalable deployment, we evaluated a budget-tiered configuration using GPT-5 Nano for high-volume evidence extraction, o4-mini for final reasoning, benchmarking against a GPT-5.4 high reasoning effort configuration applied at every step. Note-level discrepancies were adjudicated by reviewers into final co-produced labels that were used to refine labels and inform agent development. The main outcome measure was the mean absolute error (MAE) of the initial and final agent vs the final co-produced labels. The agreement between agent- and reviewer-flagged ambiguous notes, per- note cost and compute time across configurations were also tested.

Results

Expert reviewers labeled 626 notes from 273 patients; human-AI adjudication revised 127 (20%) of these initial labels and added 60 newly labeled notes, yielding a 686-note co-produced reference. Against this reference, the final agent’s accuracy improved from a mean absolute error of 0.406 to 0.291 with co-learning, and its ambiguity flag agreed with expert ambiguity designations with 92.1% accuracy. Applied across the cohort, the agent labeled 101,691 notes; the budget tiered configuration matched the accuracy of GPT-5.4 at high reasoning effort while reducing estimated cost by 69% and compute time by 70%.

Conclusion

Adopting a framework for human-AI co-learning, SHARE, improved the overall quality of gold-standard labels, identified ambiguous cases for further review, and supported accurate and standardized chart reviews of disease activity at a scale infeasible for manual review. SHARE’s resource efficiency provides a transferable approach to incorporate complex phenotypes in RWE studies.

Key messages

What is already known on this topic

Defining disease states from electronic health record (EHR) data is central to generating real-world evidence (RWE), but complex phenotypes require extensive review of narrative clinical notes that are difficult to standardize, audit, and scale. Out-of-the-box large language model (LLM) prompting can support review, but accurate annotation with face validity requires workflows that preserve supporting evidence, recognize uncertainty, while keeping the clinical experts in the adjudication loop.

What this study adds

We developed and evaluated the Synergistic Human-Agent REasoning system (SHARE), a multi-stage human-artificial intelligence (AI) co-learning framework in which clinical experts define phenotype guidelines and the agent identifies informative notes, extracts supporting evidence, assigns labels, and flags ambiguity for focused review. With rheumatoid arthritis disease activity as a use-case, adjudication improved and expanded the reference labels, while selective use of lower- and higher-cost models supported internally evaluated, resource-efficient scaling.

How this study might affect research, practice or policy

SHARE introduces a framework for human-AI workflows for research, shifting review of complex EHR phenotypes from broad manual abstraction towards a scalable, resource- efficient, targeted expert adjudication, with clinical experts defining the guidelines, overseeing local validation and the final interpretation.

Article activity feed