Reasoning Before Disposition: A Model-Agnostic Cannot-Miss Discipline for Quiet Emergencies and the Case for Deterministic Enforcement
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Large language models now match clinicians on medical-knowledge benchmarks. Whether they can safely triage a patient message is a separate question, and the highest-consequence failure in triage is the emergency that never gets escalated. We tested two competing explanations for that failure. The first is sycophancy: the model defers to a patient who downplays a danger sign. The second is atypicality: the model reads an early, mild, or atypical presentation of a time-critical disease as benign. To isolate the effect of a model-agnostic cannot-miss discipline, we performed a prompt-level ablation on a fixed frontier model (Claude Opus 4.8), then replicated the same instruction-level intervention across eight models from two families (Claude Opus 4.8, Sonnet 5, Haiku 4.5, and Fable 5; Gemini 3.1 Pro-Preview, 3.6 Flash, 3.5 Flash-Lite, and 3.1 Flash-Lite) under a documented harness and assembled the full apparatus as a model-agnostic evaluation suite.
Minimization turned out to be harmless when danger is overt. Across 40 unambiguous emergencies rendered in five framings on Opus 4.8 (neutral, mild minimization, strong denial, third-party reassurance, and a benign distractor), sensitivity stayed at or above 97.5% in every arm and was flat across framings, and seven of the eight models held 97.5% to 100% on the neutral set (the eighth completed 13 of 40 stems, all escalated). Atypicality is where the hazard lives. Across 76 subtle or atypical emergencies, each adjudicated by a blinded, LLM-simulated three-physician panel and each mapped to a published guideline that names its cannot-miss diagnosis, the base model failed to escalate 13.2% (95% CI 7.3 to 22.6) to the top acuity tier; the governance prompt cut this to 3.9% (1.4 to 11.0; paired exact McNemar 7 to 0, p = 0.016; relative risk 0.30; number needed to treat 11). The effect sits in the harder half of the cohort (36 original stems: 22.2% to 8.3%, 5 to 0; 40 extension stems: 5.0% to 0%, 2 to 0), and its magnitude holds when the three stems an independent panel rated below emergency are excluded (9.6% to 2.7%, 5 to 0, p = 0.063). A one-line distillation reproduced most of the effect, showing that the underlying reasoning principle does not depend on proprietary prompt language. Its residual failures and weaker specificity behavior also show that portability of the principle is not equivalent to deterministic enforcement. The effect replicated in Spanish (27.8% to 11.1%, McNemar 6 to 0, p = 0.031). The effect replicated in a second, documented harness on the same model, and then across eight models.
In the eight-model replication, baseline under-triage ranged from 2.6% to 34.2% by model and elicitation. The same prompt text, unchanged, was applied above eight models from two vendors, three capability tiers, and two elicitation modes; under-triage fell in every cell where there was material hazard to remove, and no cell of any model gained a false emergency alarm. The identical governance prompt reduced it in seven of twelve model-by-elicitation cells, moved a single stem in three, and was neutral or reversed in the two cells already at the floor. With the stem as the unit of analysis, governance was favored on 19 stems and control on 5 across the eight reasoning-permitted models (sign test p = 0.007); pooled discordant pairs were 27 to 8 (p = 0.002) under forced-choice elicitation in the Claude family and 31 to 10 (p = 0.001; paired risk difference 3.5 points, 95% interval 1.5 to 5.7) under reasoning-permitted elicitation across both families. The misses concentrate: 42 of the 76 stems were never missed by any model in either arm, ten stems account for 71% of all baseline misses, and a handful of conditions (subacute endocarditis, early ectopic pregnancy, early urosepsis, mesenteric ischemia, early appendicitis) are missed by most models and resist most of the prompt’s effect. On a 44-stem panel-adjudicated non-emergency cohort, the full prompt added zero escalations to EMERGENCY in any model under any elicitation (0 of 44 per cell; Wilson upper bound 8.0%), and it moved some ROUTINE stems to same-day URGENT (17 added against 2 removed across eight models, p < 0.001). Across more than 5,900 triage classifications, no emergency was ever routed home in any arm; the failure mode is uniformly EMERGENCY to URGENT. The prompt changed nothing on knowledge (MedQA 95.8% in both arms, 4 to 4 discordant), adversarial robustness, or medication-hazard detection, while raising named-guideline citation from 65.6% to 96.9% at roughly 2.3 times fewer output tokens; HealthBench fell 3.6 points on a single seed, concentrated in context-seeking and uncertainty-handling behaviors, and warrants follow-up. Elicitation mattered: forcing an immediate single-shot disposition raised baseline hazard by roughly ten percentage points in two of four Claude models, and the governance increment in the Claude family was largest under that elicitation.
Quiet emergencies, defined here as early, mild, or atypical presentations of time-critical conditions, were the principal source of under-triage in this constructed cohort rather than overt emergencies accompanied by patient minimization. A prompt-level cannot-miss discipline reduced that hazard across a heterogeneous set of models without adding false emergency escalations in the primary specificity cohort. However, residual failures remained concentrated, stochastic, dependent on model choice and elicitation, and undefined when output could not be parsed. These findings show both the portability of the reasoning discipline and the limits of implementing it through instructions alone. MedCanon is presented as a hybrid clinical runtime with a deterministic safety spine designed to convert these distributional improvements into versioned, testable invariants. The runtime itself was not evaluated in this study. The Level 2 evaluation specified here is required before claims about its clinical safety effect, determinism, or auditability can be made. Ground truth came from LLM-simulated panels triangulated against an independent cross-family panel, published guideline anchors, and an externally authored vignette set; independent board-certified human adjudication is required before any clinical-deployment claim.