A user-friendly, no-code, application for HIPAA-compliant automated analysis of tabular data at scale

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background: Clinical research often requires reviewing large volumes of unstructured electronic medical record (EMR) data, a time-consuming task demanding skilled personnel. Large language models (LLMs) like ChatGPT can efficiently analyze and summarize text, potentially accelerating chart review research. However, concerns about personal health information (PHI) leakage limit use of commercial chatbots. Secure, in-house LLMs within fire-walled environments can address these concerns. Aim: To develop a user-friendly, scalable application leveraging an in-house ChatGPT 4o-mini on Emplify Health's Azure cloud to automate processing of tabular clinical data while protecting PHI. Methods: The application imports tabular data (Excel or text), guides users through prompt engineering, and automatically submits data row-by-row to the LLM. It retrieves and tabulates results as specified by users, enabling both automated extraction of clinical parameters, and data summarization/interpretation. Results/Conclusions: Initially designed to identify cancer cases and extract related parameters from pathology reports, a fully generalized application was developed which has found utility in analyzing diverse data sources including cancer registries, cardiology CT reports, imaging narratives, and clinical notes. This tool has significantly accelerated research by reducing data retrieval time, allowing staff to focus on higher-value tasks like data analysis, while maintaining PHI protection.

Article activity feed