Machine learning combining FIT with up to 1,025 clinical variables: limited referral reduction but potential for faster diagnosis
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background
The faecal immunochemical test (FIT) is central to triaging symptomatic patients with suspected colorectal cancer (CRC) in UK primary care, yet only about one in eleven patients above the NICE 10 µg/g threshold have CRC. Existing prediction models attempting to improve on FIT have relied on conventional statistics and limited predictors.
Methods
GP-requested FITs with linked data (Jan 2017 - May 2025) were extracted from the Oxford University Hospitals (OUH) datawarehouse. Patients aged ≥18 with core bloods and 180-day CRC follow-up were included. Machine learning (ML) models were trained on up to 1,025 predictors: FIT, age, sex, blood tests and their time series slopes, diagnoses/procedures/prescriptions, deprivation, BMI, and ethnicity. Models comprised penalised logistic regression, generalised additive models (EBM, NAM, SNAM, NODE-GAM), decision tree ensembles (random forests, XGBoost), and a multilayer perceptron. Referral reduction versus FIT ≥10 µg/g was evaluated at model risk score thresholds capturing the same cancers (conservative) or same proportion of cancers (less conservative) as FIT. Potential to prioritise referred patients was assessed by examining whether positive predictive value (PPV) is very high (>30%) at any substantial sensitivity (>10%). Nested twice-repeated five-fold cross-validation provided unbiased estimates. An existing COLOFIT model was evaluated alongside.
Findings
62,219 individuals (746 CRC) were analysed; 30,862 patients (315 CRC) with high/low risk symptoms and buffered FITs formed the primary subset. At ≥10 µg/g, FIT had 91.4% sensitivity, 84.2% specificity, 5.6% PPV, and 99.9% NPV. No model reduced referrals when required to capture the same cancers as in the FIT ≥10 µg/g cohort. Generalised additive models achieved up to 18.5% referral reduction when detecting the same proportion but some different cancers as FIT ≥10 µg/g (EBM: 18.5%, NODE-GAM: 17.5%, SNAM: 17.4%, COLOFIT: 16.7%). At 30% sensitivity, EBM, NAM and NODE-GAM had average PPVs between 34.6%-35.0%, while FIT had a PPV of 14.6%.
Interpretation
Generalised additive models (GAMs) reduced referrals on average by 19% if a small proportion of the FIT-positive CRCs were substituted with originally FIT-negative CRCs by the models. No model, including COLOFIT, reduced referrals while capturing all FIT-positive cancers. Generalised additive models could detect about a third of CRCs faster, as one in three patients flagged by the models had CRC at 30% sensitivity.
Funding
EPSRC Centre for Doctoral Training in Health Data Science; National Institute for Health Research (NIHR) Oxford Biomedical Research Centre; Cancer Research UK.
Research in context
Evidence before this study
The faecal immunochemical test (FIT) is widely used in primary care to triage symptomatic patients for colorectal cancer (CRC) investigations. Combining it with other routinely collected data (such age and blood tests) in prediction models may improve its performance. We used existing systematic reviews to understand the performance of FIT, and to identify CRC prediction models incorporating FIT in symptomatic primary care patients. We also searched PubMed and Google Scholar for machine learning models that incorporate FIT. We inferred that only about one in eleven symptomatic primary care patients have CRC, suggesting there is potential to improve precision and better target colonoscopy resources. Only three studies had developed models specifically in the target patient population, using traditional statistical methods with at most six predictor variables, indicating that machine learning methods and more diverse predictor variables remain unexplored in the target population.
Added value of this study
Using one of the largest samples to date (62,219 symptomatic primary care patients), we applied a range of interpretable and flexible machine learning models with up to 1,025 predictor variables (including blood test trends and medical history). We evaluated the models with clinically relevant performance metrics: potential to reduce colonoscopy referrals and potential to fast-track patients for investigations. No machine learning model or an existing COLOFIT model reduced referrals when required to detect all cancers already detected by FIT at the standard 10 µg/g threshold; the models reduced referrals at most by 19% if some missed FIT-positive cancers were substituted with extra FIT-negative cancers. However, when required to detect only 30% of cancers, generalised additive models had high positive predictive values (close to 35%), which were substantially higher than the PPV of FIT at the same sensitivity (14.6%). The models may thus be useful for fast-tracking higher-risk patients for colonoscopy to detect a subset of cancers faster.
Implications of all the available evidence
FIT is a highly sensitive and specific test for CRC: it is hard to improve on its performance with routinely collected data for reducing referrals in FIT-positive patients. However, prediction models show greater potential for prioritising referrals: detecting colorectal cancers faster in a subset of higher-risk patients.