Vision and Language Models for Classifying Maxillary Sinus Disease on Cone-Beam Computed Tomography: A Transparent Multimodal Benchmark

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Cone-beam computed tomography (CBCT) frequently captures the maxillary sinuses incidentally, and reliable automated detection of sinus abnormality is clinically relevant. Unlike most vision–language benchmarks in medical imaging, which pair images with pre-existing, human-authored clinical reports, findings text can also be generated directly by a large language model from the image itself—raising the question of how much diagnostic value such AI-derived text carries, and whether that value depends on independent verification. Multimodal artificial intelligence (AI) benchmarks risk overstating performance if the provenance of each input—image, raw AI-generated text, or radiologist-verified text—is not clearly separated and reported.

Methods

We used 300 mid-sagittal CBCT slices from the MMDental dataset. ChatGPT generated findings text and a provisional normal/abnormal label for every slice (majority vote, three independent readings from the image alone); primary classification performance was assessed on this full, unfiltered set (n=300). A radiologist then independently reviewed each case’s image together with ChatGPT’s description, producing their own diagnosis; three cases were excluded as insufficient, yielding 297 confirmed cases. On this subset, every model was retrained and re-evaluated under identical 10-fold cross-validation on both the provisional ChatGPT-only labels (“pre”) and the radiologist-confirmed labels (“post”), isolating the effect of label provenance from image or architecture. Eight vision architectures, seven language classifiers, and five VLMs were evaluated throughout; three generative models performed exploratory note-drafting.

Findings

Raw ChatGPT-generated text produced the highest performance of any modality or condition: language models reached near-ceiling AUC (0·992 to 1·000, n=300), exceeding every vision model (AUC 0·799 to 0·880) and every VLM image-only probe (AUC 0·63 to 0·69). On the 297-case pre/post analysis, this advantage depended heavily on label source: language and text-derived VLM performance fell substantially from ChatGPT-only to radiologist-confirmed labels (e.g. BERT-base AUC 0·999 to 0·837), while vision-model performance was stable or modestly improved (e.g. DenseNet-121 0·867 to 0·891). The radiologist reclassified 62 of 297 cases (21%) relative to ChatGPT’s provisional read, and a meaningful proportion of raw ChatGPT text was clinically uninterpretable or unsupported by the imaging.

Interpretation

As shown here for the first time, raw, image-derived AI-generated text yields the highest apparent classification performance in this benchmark, but this reflects the text’s alignment with its own self-generated labels rather than verified diagnostic content, and a substantial share of that text is not clinically explainable. Radiologist-confirmed text and labels give a lower but trustworthy estimate of true performance, on which convolutional neural network (CNN) vision models remain a stable, comparatively inexpensive baseline. Multimodal dental AI should report performance separately by modality and label provenance rather than pooling headline metrics.

Research in context

Evidence before this study

We searched PubMed with the terms “maxillary sinus”, “CBCT”, “deep learning”, “artificial intelligence”, and “vision–language model” from January 2015 to July 2026. Published benchmarks have predominantly evaluated single-modality convolutional networks for dental CBCT tasks, with limited transparency about how AI-generated text inputs are produced or verified. No study had systematically compared image-only, language-only, and VLM families for maxillary sinus classification on the same dataset under matched cross-validation, nor had the extent to which apparent multimodal performance depends on whether AI-generated text is used raw or independently verified been explicitly characterised.

Added value of this study

Unlike prior vision–language benchmarks in medical imaging, which typically pair images with pre-existing, human-authored clinical reports, this study evaluates text generated directly by an LLM from the image itself, independent of any prior radiological report. This design isolates a question distinct from conventional report-grounded VLM evaluation: not how well a model can classify existing clinical documentation, but how much apparent diagnostic value an LLM’s own image-derived description carries—and how much of that apparent value depends on whether the description is used raw or independently verified. We show that raw, unverified ChatGPT-generated text produces the highest classification performance of any modality evaluated (AUC up to 1·000), exceeding image-only and image-plus-text models—but that a meaningful proportion of this raw text is clinically uninterpretable, and that its apparent advantage collapses substantially once evaluated against a radiologist-confirmed reference standard on the same cases. Trained vision models, by contrast, remain stable regardless of label source, indicating that their signal derives from image content rather than label provenance. We further show that ViTs underperform CNNs at dental-imaging dataset scale.

Implications of all the available evidence

Raw AI-generated text can substantially outperform image-based classification on headline accuracy metrics, but this performance is not synonymous with clinical trustworthiness or explainability and should not be reported or deployed without independent verification. CNNs remain reliable, comparatively low-cost baselines for sinus CBCT screening at small dataset sizes. Multimodal dental AI benchmarks must report modality-specific results transparently and distinguish raw from verified text inputs; headline metrics should not be pooled across these conditions without disclosing their provenance. Generative VLMs show near-term utility as documentation aids rather than autonomous diagnostic systems, and full three-dimensional CBCT analysis remains the essential next step for clinical translation in the maxillofacial trauma setting.

Article activity feed