Quantifying the Quality of Corrective Actions in Medical Safety Incident Reports Using a Hybrid Rule-Based and Large-Language-Model Classification System: A Cross-Sectional Pilot Feasibility Analysis of 11,507 Japanese National Reports (2010, Interim)

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Whether corrective actions documented in medical safety incident reports rely on individual vigilance (“Safety-I”) or on structural, system-level intervention (“Safety-II”) has not been quantitatively evaluated on a national scale in Japan. We developed an automated classification pipeline to assign corrective-action free-text to a 7-level maturity scale (L0–L6) and computed two summary indices: the Safety Measure Quality Profile (SMQP), the full L0–L6 distribution, and the System-based Safety Measure Rate (SSMR), the proportion of non-L0 records classified L3–L6.

Methods

We analyzed all 11,507 corrective-action free-text entries from the 2010 release of Japan’s national medical accident and near-miss reporting database (Japan Council for Quality Health Care, JCQHC), comprising 8,804 near-miss (Hiyari-Hatto) and 2,703 accident (Jiko) reports. Records were classified using a five-stage hybrid pipeline: an expert-developed rule dictionary, TF-IDF + k-nearest-neighbor matching, cosine-similarity matching, a two-tier large-language-model (LLM) classifier, and a conservative priority-cascade fallback. SSMR was compared between near-miss and accident reports using a χ 2 test, Wilson 95% confidence intervals, Cramér’s V, and the risk difference (RD), against pre-specified minimal clinically important difference (MCID) criteria of RD ≥ 2 percentage points and Cramér’s V ≥ 0.10.

Results

Every record received a definitive L0–L6 label (0% unresolved). Overall, 16.6% of records were unclassifiable (L0); among the 9,599 classifiable (non-L0) records, individual-vigilance actions (L1) predominated (54.6% of all records), and only 11.82% (95% CI, 11.19–12.49%) met the SSMR criterion (L3–L6). SSMR was higher for accident reports than for near-miss reports (18.12% [95% CI, 16.69–19.64%] vs. 9.46% [95% CI, 8.80–10.17%]; RD = 8.66 percentage points; Cramér’s V = 0.120; χ 2 (1) = 136.97, p <0.001), exceeding both pre-specified MCID thresholds.

Conclusions

In this interim single-year analysis, the large majority of documented corrective actions in Japanese medical safety reports remained individual-vigilance-based rather than system-based, with accident reports showing a substantively, rather than merely statistically, higher proportion of system-based actions than near-miss reports. These findings support the feasibility of large-scale automated assessment of corrective-action quality and provide the rationale for the planned 16-year longitudinal analysis.

Status note

This report presents an interim, single-year (2010) result generated during the rule-dictionary development and validation phase of a planned longitudinal study spanning 2010–2025 (16 years, N = 178,375; UMIN000071271). This pilot report corresponds to Study 2 (SMQP/SSMR: corrective-action quality assessment) of the planned longitudinal study and is intended to evaluate methodological feasibility before application to the complete 2010–2025 corpus. Findings reported here should not be interpreted as the study’s primary longitudinal outcome and are subject to revision once the classification dictionary is frozen and applied to the full corpus.

Article activity feed