Comparative Evaluation of a System One Model and a General-Purpose Large Language Model on the Korean Physical Therapist Licensing Examination
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background: System One Model is designed for structured decision-making and can provide probabilistic outputs, but their performance in domain-specific physical therapy tasks has not been established. Objective: To evaluate the performance and potential utility of Jev, a System One Model, for professional knowledge-based selection tasks in physical therapy by comparing it with GPT-4o. Methods: A total of 380 multiple-choice questions from the 2024 and 2025 Korean Physical Therapist Licensing Examinations were independently submitted to Jev and GPT-4o. The outcomes included answer accuracy, subgroup performance, output validity, API response latency, token usage, estimated API cost, and Jev's prediction uncertainty. A retrospective probability-based model-cascading simulation was also performed. Results: Jev correctly answered 282 questions (74.2%), whereas GPT-4o answered 326 (85.8%), a difference of -11.6 percentage points. Jev produced valid structured responses for all questions, whereas GPT-4o produced three invalid responses. Jev's selected-answer probabilities discriminated correct from incorrect answers (AUROC = 0.866), and accuracy reached 99.3% when probabilities were greater than or equal to 0.90. The median API response latency was 0.94 s for Jev and 1.06 s for GPT-4o, with estimated costs of US$0.008 and US$0.161, respectively. At a retrospective cascading threshold of 0.70, accuracy was 86.6% with 175 simulated GPT-4o calls. Conclusion: Jev showed lower overall accuracy than GPT-4o but provided reliable structured outputs, informative prediction probabilities, lower latency, and lower estimated cost. Its probabilistic outputs warrant further prospective evaluation for selective model escalation and decision-support workflows in physical therapy.