Auditing the Human–LLM Autonomy Gap in Clinical Ethics: Development and Application of the Autonomy Index Across 50 Clinical Ethics Vignettes

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background.

Respect for autonomy is central to biomedical ethics but is rarely assessed with a structured, reproducible measure. This gap matters as patients, families, and clinicians increasingly use general-purpose language models for health-related questions. We developed the Autonomy Index (AIx) to characterize autonomous agency in clinical-ethics vignettes and compared ratings from trained human reviewers with ratings from current language models.

Methods.

We reviewed ten decisional-capacity instruments and classified them by target construct into consent capacity, decision capacity, and functional autonomy, identifying a region of autonomous agency (values, deliberation, and enactment) that existing instruments measure least well. The resulting thirteen-item instrument scores four domains: Values Awareness, Factual Understanding, Rational Deliberation, and Intentional Action, on a five-point ordinal scale with an explicit not-applicable option; the External Constraint Index and Support Provided Index are reported separately. Nine models scored 89 clinical vignettes. We selected 50 vignettes by model-score tier and inter-model disagreement, then obtained 173 ratings from 30 clinical and bioethics-informed reviewers and 3,431 valid scores from eight models in July 2026. The prespecified comparison summarized each source within vignette and estimated the paired human-minus-model difference; mixed-effects, domain, agreement, and variance-component analyses were secondary or exploratory. Human data were collected under Harvard Longwood Campus IRB protocol IRB26-0146.

Results.

Across the selected 50-vignette comparison set, human and model composites were strongly associated (Pearson r = 0.781, P = 2.3 × 10⁻¹¹; Spearman ρ = 0.775). The human mean was 50.2 and the model mean was 39.2, a paired human-minus-model difference of 11.0 index points (95% CI, 6.3 to 15.8; P = 2.6 × 10⁻⁵; d_z = 0.66). Human means exceeded model means on 36 of 50 vignettes. Bland–Altman limits of agreement were wide (−21.9 to 43.9), indicating substantial case-level variation. Mean differences were positive in all four domains: Factual Understanding, 13.7; Intentional Action, 12.3; Values Awareness, 10.4; and Rational Deliberation, 7.9 points; the joint domain-by-source test did not detect heterogeneity (Wald χ²[3] = 2.27, P = 0.52). Vignette content accounted for 80.9% of score variance, whereas source accounted for 0.7%. Models never selected not applicable; human reviewers did so for 22.5% of item responses. Among 19 vignettes with at least four human ratings, 9 had human-minus-model differences greater than 15 points and none had a difference below −15 points.

Conclusions.

In this deliberately selected 50-vignette set, trained reviewers assigned higher autonomy scores than the tested general-purpose language models on average, while the two sources ranked cases similarly. Wide limits of agreement and variation across well-rated vignettes preclude treating the mean difference as a prediction for an individual case. Because the comparison set was selected using prior model scores, reviewers were a convenience sample, and human ratings are a reference standard rather than ground truth, the findings should be interpreted as evidence of systematic divergence in this sample, not proof that models underestimate patient capacity. External validation in prospectively sampled cases with denser human rating is required before clinical use.

Article activity feed