Comparative Analysis of Clinical Finding Retrieval Difficulty in Chest X-ray Retrieval Models

This article has been Reviewed by the following groups

Read the full article

Abstract

Automated chest X-ray (CXR) report generation increasingly relies on retrieval-based systems that select sentences from fixed candidate pools of clinical language. Because these pools may be unevenly distributed across clinical findings, this study investigated whether candidate-pool composition affects finding-level retrieval performance and cross-model agreement on relative finding difficulty. Three sentence-level image-to-text retrieval models, a proposed model based on the JoImTeRNet framework, CXR-RePaiR, and CXR-ReDonE, were evaluated using finding-level recall, precision, and F1 at K=1, 5, and 10 under two candidate-pool conditions. The original pool contained 129,906 sentences from the MIMIC-CXR findings sections, while a balanced, reduced-size candidate pool capped positive sentences per finding at 450, yielding 6,216 sentences. Cross-model agreement in finding difficulty was assessed using Spearman rank correlation with paired bootstrap testing. Overall recall ranged from 0.09 to 0.31 and precision from 0.15 to 0.29 across models and K values. Lung Opacity and Pleural Effusion were consistently among the most difficult findings to retrieve, whereas Support Devices was consistently among the easiest. Under the original pool, cross-model agreement on finding difficulty was strong (ρ=0.749–0.807) but decreased markedly under the balanced pool (ρ=0.240–0.495), with significant reductions for all model pairs (∆ρ=0.298–0.487, all p<0.001). These findings indicate that candidate-pool composition substantially influences both finding-level retrieval performance and cross-model agreement, indicating that apparent consensus on retrieval difficulty may partly reflect evaluation-pool composition rather than intrinsic model behavior.

Article activity feed

  1. This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/22788887.

    In order to evaluate the impact of the candidate-sentence pool on the retrieval of clinical findings for chest X-ray images, we investigate three different sentence-retrieval models, namely the proposed JoImTeRNet model as well as CXR-RePaiR and CXR-ReDonE. We assess these models under two candidate-sentence pool conditions: an original, unbalanced pool of 129,906 sentences and a balanced pool of 6,216 sentences, capped at 450 positive sentences per finding. Our results show that candidate-pool composition strongly affects both finding-level retrieval performance and cross-model agreement on which findings are "hard" to retrieve. Cross-model agreement was strong for the original pool of sentences (ρ=0.749–0.807), but dropped for the balanced pool of sentences (ρ=0.240–0.495). Lung Opacity and Pleural Effusion were consistently the two hardest findings to retrieve, whereas Support Devices was consistently the easiest.

    Field contribution

    It is novel to empirically demonstrate that the apparent "consensus" between independently trained retrieval models in a multi-model setup does not necessarily stem from the models themselves, but rather may be influenced by the design of the evaluation pool.

    Major issues

    • Confounded variables: The authors note that the "balanced" pool differs from the original pool in terms of size, class balance, and even the included sentences. However, rather than subsampling to a fixed size, the authors use a substantially smaller balanced pool. An equal-size random subsample would be needed to isolate the effect of balancing from the effect of pool size.

    • Possible bias favoring the proposed model: The sentence pool for evaluation was selected from the MIMIC-CXR database used to train the proposed model, while CXR-ReDonE was trained on the stylistically different MIMIC-PRO. This could confer an advantage unrelated to genuine retrieval quality.

    • Evaluation labels not expert-verified: The evaluation set is a random sample from the CheXpert training set, not the expert-adjudicated validation set. Moreover, the ground truth was generated by an automated labeler (CheXbert), hence introducing label noise.

    • Small number of findings (14) for correlation analysis: With Spearman's rank correlation based on only 14 findings, estimates will inherently be unstable, even with the use of bootstrap confidence intervals.

    • Speculative explanation for Lung Opacity's difficulty: The hypothesis in the Figure that the findings' semantic/definitional ambiguity contributes to the difficulty of retrieving Lung Opacity is plausible but was never directly tested.

    Minor issues

    • Some redundancy between the abstract, results summary, and conclusion.

    • The "Related Work" section is rather long in comparison to the content of the paper; it could be shortened.

    • Some of the tables, especially Table 3, contain a lot of numbers and therefore would benefit from a visual summary alongside Figure 1.

    • Justification for the specific number of 450 positive sentences for the balancing cap is not clear.

    • Reference [21] is listed as a source but is not used in the paper.

    Competing interests

    The author declares that they have no competing interests.

    Use of Artificial Intelligence (AI)

    The author declares that they did not use generative AI to come up with new ideas for their review.