From Output Errors to Workflow Harm: A Practitioner-Audit Method for LLM-Mediated Research
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Objective
Formal large language model (LLM) evaluations score isolated prompts, but clinicians and health-informatics researchers meet model failures inside multi-step workflows where erroneous output can alter procedures or contaminate documents. We present TRACE (Tracking Reliability of AI-generated Conversational Evidence), a practitioner-audit framework for evaluating the downstream workflow reliability of conversational AI.
Materials and Methods
A method paper with an empirical demonstration: 45 documentation-positive incidents recorded by one clinician-informatician across scholarly, clinical informatics, and clinical-adjacent workflows over seven weeks, coded with a consequence-based severity rubric, an error definition, a taxonomy crosswalk, and a Response-Audit Scorecard. Three reviewer-authors independently coded a 16-incident subsample; three vendor-blinded AI comparators applied the taxonomy to all 45 incidents.
Results
Four categories tied as most frequent: verification failure, factual numerical error, tool-behavior misunderstanding, and citation or reference formatting (n=7 each). Four workflow-harm patterns recurred— procedural propagation, documentary contamination, trust-calibration disruption, and user-borne corrective burden—and one incident carried an estimated $2500 impact. Category agreement across three human reviewer-authors was low (Fleiss κ=0.155), whereas three AI comparators agreed substantially (Fleiss κ=0.632), suggesting taxonomy legibility under standardized conditions even where human judgment diverged.
Discussion
Category assignment is comparatively legible, whereas severity and claimed-verification remain judgment-dependent. The claimed-verification gap is a measurable failure mode distinct from hallucination, sycophancy, and over-refusal.
Conclusion
Practitioner audits with structured response scoring complement benchmarks by documenting workflow harm as an applied evaluation unit for clinical informatics and public-health work; this is a pilot that motivates, not estimates, error rates or cross-model comparisons.