Cross-System Legibility of a Practitioner-Derived Workflow-Error Taxonomy for Conversational AI: A Three-Comparator Agreement Study
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Practitioner-derived taxonomies of conversational artificial intelligence (AI) workflow errors show low inter-rater agreement among human coders, leaving open whether the instrument is ill-specified or the judgments are inherently difficult. We delivered a locked eight-category workflow-error taxonomy verbatim, under standardized conditions, to three frontier large language model comparators from distinct developer lineages, which coded a documented 45-incident error corpus. On the same 16 incidents coded by three human reviewer-authors, comparator category agreement was substantial (Fleiss κ=0.625) against slight human agreement (κ=0.155); across the full corpus it was stable (κ=0.632), and a 10-category refinement did not reduce it (κ=0.690). Severity and a claimed-verification flag remained only fair in both arms and on both samples. Substantial cross-system consistency provides a legibility signal consistent with recoverable category distinctions, but cannot separate instrument clarity from shared model priors, and is not a validation of any coding.