Optical compression of multimodal clinical data into a unified visual memory for generative trajectory forecasting
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Importance
Current clinical artificial intelligence is bottlenecked by highly specialized, fragmented models that reduce complex, multimodal data into isolated scalar predictions. By reframing multimodal perception as a pure vision problem, optical compression into a universal visual operating system can resolve internal data interaction bottlenecks and shift the paradigm toward continuous, generative clinical forecasting.
Objective
To develop and internally validate Clinical Visual Memory, a generalist visual foundation model utilizing high-fidelity optical compression to fuse five heterogeneous intensive care unit (ICU) data modalities into a unified visual-token space, acting as a generative navigation system for critical-care trajectories.
Design, Setting, and Participants
Retrospective modeling study using MIMIC-IV adult ICU encounters. After excluding stays shorter than 24 hours, 54,551 patients, 68,546 hospital admissions, and 74,829 ICU stays comprised the eligible cohort. Strict patient-level partitioning prevented cross-split leakage.
Methods
Structured electronic health record (EHR) data, vital signs, 10-second electrocardiogram (ECG) waveforms, chest radiographs, and clinical notes were rendered as 2D images and encoded by a single frozen DINOv2 vision transformer. Through modality-aware latent-query cross-attention, these highly heterogeneous sources were optically compressed into a shared 1024-dimensional Clinical Visual Memory. To establish a universal output interface, this memory conditioned an conditional instruction-tuned image generator to decode eight-domain deterioration trajectories across 3- to 48-hour horizons. Visual compression fidelity was evaluated by reducing retained source-pixel area to 1%, and critical transitions were mapped using event-specific projected-axis geometry.
Results
Operating as a generalist foundation, Clinical Visual Memory achieved AUROCs of 0.852 (95% CI, 0.821-0.882) for 48-hour mortality, 0.699 (0.673-0.725) for incident acute kidney injury (AKI), 0.723 (0.688-0.757) for high Sequential Organ Failure Assessment (SOFA), and 0.742 (0.726-0.757) for alive ICU discharge. Crucially, resolving the token bottleneck via an 8-fold reduction in source-pixel area preserved 98.9% of the uncompressed 48-hour mortality AUROC. Generative image-out decoding maintained clinical meaning, and trajectory geometry functioned as a proactive navigation system, yielding median warning lead times of 34.8 hours (IQR, 17.5-43.3) before death, 14.6 hours (6.5-33.2) before AKI, 15.0 hours (6.0-27.0) before high SOFA, and 20.8 hours (10.2-36.0) before recovery, with low false-alert burdens (0.040-0.109 per patient-day).
Conclusions and Relevance
Clinical Visual Memory demonstrates that optical compression can successfully unify highly heterogeneous clinical data into a single, high-fidelity visual-token representation. By functioning as a continuous clinical navigation system rather than a discrete alert generator, this framework lays the architectural foundation for a generalist, generative visual operating system in critical care medicine. External and prospective validation are required before clinical deployment.
KEY POINTS
Question
Can optical compression of heterogeneous, multimodal intensive care data into a unified visual-token memory enable a generalist generative framework for continuous clinical forecasting?
Findings
In MIMIC-IV, five clinical modalities were rendered as images and fused via a shared frozen DINOv2 transformer into a 1024-dimensional Clinical Visual Memory. This unified architecture supported robust multitask prediction, achieving AUROCs of 0.852 for 48-hour mortality, 0.699 for incident AKI, 0.723 for high SOFA, and 0.742 for alive ICU discharge. Crucially, an 8-fold reduction in source-pixel area preserved 98.9% of the uncompressed mortality discrimination, validating high-fidelity optical compression. Generative image-out decoding successfully translated this memory into actionable clinical trajectories, yielding median warning lead times of 34.8 hours before death, 14.6 hours before AKI, 15.0 hours before high SOFA, and 20.8 hours before alive ICU discharge.
Meaning
By resolving multimodal data bottlenecks through high-fidelity visual compression, this framework shifts clinical AI from fragmented, task-specific alerts to a unified, generalist visual operating system. It demonstrates that complex patient states can be continuously navigated and generatively forecasted via a single interpretable visual-token memory.