From National Data to Local Evaluation: Benchmarking Machine Learning Models for Perioperative Risk Prediction

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Whether newer modeling approaches improve perioperative risk prediction is unclear. Holding 73 prespecified preoperative variables fixed (69 observed and used as tabular inputs, all 73 serialized to text), we benchmarked six tabular model families and three pretrained clinical text-encoder embedding classifiers for four 30-day outcomes, training on 4,995,670 ACS NSQIP cases (2018-2022) and testing temporally on 963,565 cases (2024). The FT-Transformer led or tied on AUROC for all four outcomes (0.755-0.955), exceeding logistic regression by only 0.009-0.030; the three encoders fell 0.005-0.009 short of the best tabular model and spanned 0.003 or less. On identical cases the ACS NSQIP Surgical Risk Calculator matched the best model for mortality but trailed for any complication. Rankings changed under AUPRC and did not hold in a 3,490-case single-institution deployment check, not independent of the PUF. Once the predictor set is fixed, architecture and encoder choice move discrimination little.

Article activity feed