Automated Skill Optimisation for False Presupposition Handling in Cancer Communication: SkillOpt Versus Conventional Prompt Engineering Across Language Models

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background: Large language models (LLMs) may answer cancer questions plausibly while accepting a false presupposition (FP). Prompts that increase FP correction may also challenge questions in which no FP is present (NFP). We compared SkillOpt, an automated skill-optimisation method developed by Microsoft Research, with conventional prompt engineering across models. Methods: Using the physician-verified Cancer-Myth benchmark, GPT-5.4 optimised a written skill for GPT-4o in three independent 12-step executions. Training used 117 FP questions. Rewrites were selected by combined accuracy on a validation set of 58 FP and 30 NFP questions, then tested on 410 unseen FP questions. A separate GPT-4o judged responses. Across seven LLMs, we compared no correction instruction, a deliberately strong seven-example few-shot prompt, three GPT-5.4-generated instruction-only prompts, and full and compact SkillOpt skills. Further tests examined portability, new questions and valid clinical assumptions. Results: SkillOpt increased GPT-4o FP correction from 15.4-17.6% to 82.2-90.5%, gains of 66.3-72.9 percentage points (95% bootstrap confidence intervals 61.7-77.3; all Holm-adjusted p<0.001). Only two or three rewrites were accepted per execution. The first full skill achieved 92.7-96.6% correction on three recipient models but 6.3-16.3% on two weaker models, where few-shot prompting achieved 44.6-54.6%. On Gemini 3.5 Flash and GPT-5.6 Luna, compact SkillOpt retained 90.0-92.9% correction while increasing NFP accuracy from 59.2-79.2% with the full skill to 85.8-93.3%. Both approaches sometimes challenged valid clinical assumptions. Conclusions: SkillOpt produced large, repeatable gains, but it was neither universally portable nor free of over-correction. Automated and conventional prompts should be compared for each intended model using tests of both FP correction and NFP accuracy. Compact SkillOpt may improve this balance on capable models.

Article activity feed