Automated Skill Optimisation for False Presupposition Handling in Cancer Communication: SkillOpt Versus Conventional Prompt Engineering Across Language Models
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background: Large language models (LLMs) may answer cancer questions plausibly while accepting a false presupposition (FP). Prompts that increase FP correction may also challenge questions in which no FP is present (NFP). We compared SkillOpt, an automated skill-optimisation method developed by Microsoft Research, with conventional prompt engineering across models. Methods: Using the physician-verified Cancer-Myth benchmark, GPT-5.4 optimised a written skill for GPT-4o in three independent 12-step executions. Training used 117 FP questions. Rewrites were selected by combined accuracy on a validation set of 58 FP and 30 NFP questions, then tested on 410 unseen FP questions. A separate GPT-4o judged responses. Across seven LLMs, we compared no correction instruction, a deliberately strong seven-example few-shot prompt, three GPT-5.4-generated instruction-only prompts, and full and compact SkillOpt skills. Further tests examined portability, new questions and valid clinical assumptions. Results: SkillOpt increased GPT-4o FP correction from 15.4-17.6% to 82.2-90.5%, gains of 66.3-72.9 percentage points (95% bootstrap confidence intervals 61.7-77.3; all Holm-adjusted p<0.001). Only two or three rewrites were accepted per execution. The first full skill achieved 92.7-96.6% correction on three recipient models but 6.3-16.3% on two weaker models, where few-shot prompting achieved 44.6-54.6%. On Gemini 3.5 Flash and GPT-5.6 Luna, compact SkillOpt retained 90.0-92.9% correction while increasing NFP accuracy from 59.2-79.2% with the full skill to 85.8-93.3%. Both approaches sometimes challenged valid clinical assumptions. Conclusions: SkillOpt produced large, repeatable gains, but it was neither universally portable nor free of over-correction. Automated and conventional prompts should be compared for each intended model using tests of both FP correction and NFP accuracy. Compact SkillOpt may improve this balance on capable models.