Benchmarking modular skills and fixed orchestration across six large language models in a meta-analysis task
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background. Large language models (LLMs) are increasingly used to support systematic review and meta-analysis, but complex evidence flow, trial-family identification, study-level data extraction, and statistical execution remain vulnerable to omissions, untraceable data paths, and scientifically flawed steps. It remains unclear whether gains are driven by modular skills or by fixed orchestration, and whether these effects differ across base models. Methods. We conducted a controlled, multi-model benchmark in which six LLMs completed the same randomized controlled trial (RCT) meta-analysis task under three conditions: (A) no meta-analysis-specific skills; (B) autonomous access to modular meta-analysis skills; and (C) the same skills executed under a Quality-First fixed directed acyclic graph (DAG). Each model was run three times per condition (54 runs total). The primary outcome was the scientific performance score (SPS, 0-100), rated by two meta-analysis experts; one expert was masked to explicit model and condition labels during SPS scoring, and final scores were the arithmetic mean of the two raters. Inter-rater reliability was high (SPS ICC(A,2) = 0.991). Results. Mean SPS was 43.50 (condition A), 55.65 (B), and 54.44 (C). Relative to unassisted models, modular skills improved mean SPS by +12.15 points (unadjusted 95% CI 3.15-21.16; Holm-adjusted P = 0.033). The combined modular-skills-plus-fixed-DAG condition showed a positive but not statistically significant difference relative to unassisted models (C-A = +10.94; Holm-adjusted P = 0.073) and no incremental gain over skills alone (C-B = -1.21; Holm-adjusted P = 0.746). No condition-by-model interaction was detected (P = 0.801). Conclusions. In this benchmark task, access to modular skills improved mean scientific performance relative to unassisted execution; fixed orchestration of the same skills did not provide a detectable incremental benefit. The findings are task-specific and do not establish general effects for other meta-analysis settings, models, or orchestration designs. Keywords: Large language models; meta-analysis; Agent skills; Workflow orchestration; Evidence synthesis