Cross-System Legibility of a Practitioner-Derived Workflow-Error Taxonomy for Conversational AI: A Three-Comparator Agreement Study

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Practitioner-derived taxonomies of conversational artificial intelligence (AI) workflow errors show low inter-rater agreement among human coders, leaving open whether the instrument is ill-specified or the judgments are inherently difficult. We delivered a locked eight-category workflow-error taxonomy verbatim, under standardized conditions, to three frontier large language model comparators from distinct developer lineages, which coded a documented 45-incident error corpus. On the same 16 incidents coded by three human reviewer-authors, comparator category agreement was substantial (Fleiss κ=0.625) against slight human agreement (κ=0.155); across the full corpus it was stable (κ=0.632), and a 10-category refinement did not reduce it (κ=0.690). Severity and a claimed-verification flag remained only fair in both arms and on both samples. Substantial cross-system consistency provides a legibility signal consistent with recoverable category distinctions, but cannot separate instrument clarity from shared model priors, and is not a validation of any coding.

Article activity feed