Comparison of Deep Learning Approaches for Extreme Low-SNR Image Restoration
This article has been Reviewed by the following groups
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
- Evaluated articles (GigaScience)
Abstract
Background
Live-cell fluorescence microscopy enables the study of dynamic cellular processes. However, fluorescence microscopy can damage cells and disrupt these dynamic processes through photobleaching and phototoxicity. Reducing light exposure mitigates the effects of photobleaching and phototoxicity but results in low signal-to-noise ratio (SNR) images. Deep learning provides a solution for restoring these low-SNR images. However, these deep learning methods require large, representative datasets for training, testing, and benchmarking, as well as substantial GPU memory, particularly for denoising large images.
Results
We present a new fluorescence microscopy dataset designed to expand the range of imaging conditions and specimens currently available for evaluating denoising methods. The dataset contains 324 paired high/low-SNR images ranging from four to 282 megapixels across 12 sub-datasets that vary in specimen, objective used, staining type, excitation wavelength, and exposure time. The dataset also includes spinning disk confocal microscopy examples and extreme-noise cases. We evaluated three state-of-the-art deep learning denoising models on the dataset: a supervised transformer-based model, a supervised CNN model, and an unsupervised single image model. We also developed an image stitching method that enables large images to be processed in smaller crops and reconstructed.
Conclusions
Our dataset provides a diverse benchmark for evaluating deep learning denoising methods, and our stitching method provides a solution to GPU memory constraints encountered when processing large images. Among the evaluated deep learning models, the supervised transformer-based model had the highest denoising performance but required the longest training time.
Article activity feed
-
AbstractBackground Live-cell fluorescence microscopy enables the study of dynamic cellular processes. However, fluorescence microscopy can damage cells and disrupt these dynamic processes through photobleaching and phototoxicity. Reducing light exposure mitigates the effects of photobleaching and phototoxicity but results in low signal-to-noise ratio (SNR) images. Deep learning provides a solution for restoring these low-SNR images. However, these deep learning methods require large, representative datasets for training, testing, and benchmarking, as well as substantial GPU memory, particularly for denoising large images.Results We present a new fluorescence microscopy dataset designed to expand the range of imaging conditions and specimens currently available for evaluating denoising methods. The dataset contains 324 paired …
AbstractBackground Live-cell fluorescence microscopy enables the study of dynamic cellular processes. However, fluorescence microscopy can damage cells and disrupt these dynamic processes through photobleaching and phototoxicity. Reducing light exposure mitigates the effects of photobleaching and phototoxicity but results in low signal-to-noise ratio (SNR) images. Deep learning provides a solution for restoring these low-SNR images. However, these deep learning methods require large, representative datasets for training, testing, and benchmarking, as well as substantial GPU memory, particularly for denoising large images.Results We present a new fluorescence microscopy dataset designed to expand the range of imaging conditions and specimens currently available for evaluating denoising methods. The dataset contains 324 paired high/low-SNR images ranging from four to 282 megapixels across 12 sub-datasets that vary in specimen, objective used, staining type, excitation wavelength, and exposure time. The dataset also includes spinning disk confocal microscopy examples and extreme-noise cases. We evaluated three state-of-the-art deep learning denoising models on the dataset: a supervised transformer-based model, a supervised CNN model, and an unsupervised single image model. We also developed an image stitching method that enables large images to be processed in smaller crops and reconstructed.Conclusions Our dataset provides a diverse benchmark for evaluating deep learning denoising methods, and our stitching method provides a solution to GPU memory constraints encountered when processing large images. Among the evaluated deep learning models, the supervised transformer-based model had the highest denoising performance but required the longest training time.
This work has been peer reviewed in GigaScience(see https://doi.org/10.1093/gigascience/giag071), which carries out single-anonymized peer review. These reviews are published under a CC-BY 4.0 license and were as follows:
Reviewer 2:
In the manuscript, the authors present a dataset of 324 paired high- and low-signal-to-noise ratio fluorescence microscopy images designed to improve the training and benchmarking of deep learning denoising models across various biological specimens and imaging conditions. The authors also developed an image stitching method to address GPU memory constraints when processing large images and demonstrated that supervised transformer-based models achieve the highest denoising performance among state-of-the-art methods. Overall I found the exposition quite good but some parts could result a bit confusing. So while I think the work is useful, I have some comments.
Here are my points:
I think the way Table 1 is organized is not very clear (at least to me!). Because the table refers to images with different sizes I was quite a bit lost. If for technique A, Sample B, there is an image of 26 MP (which I guess stands for megapixels), is this image then divided in 100 non-overlapping 512x512 images? So of these 100 images 90 are considered for training? How does this really work? In the table, instead of the MP indication, I would put how many paired images are considered for training/testing/validation and their typical size (e.g. technique A, Sample B has 100 images for training, 20 for testing, and 10 for validation. Also I would mention that these images have all size 512x512 or whatever the size was. I actually found a hint of this in the methods section toward the end of the paper. But I would just insert explicit numbers in the table so the reader knows immediately what is happening from the beginning.
The term "high-resolution images" is used ambiguously ("We introduce a novel dataset of 324 high-resolution images"). Does this refer to high pixel counts (large field of view), high spatial resolution (sampling frequency/Nyquist), or the optical resolution of the objectives used? A clearer definition is required.
Given the difference in image sizes across the dataset, some samples appear to contribute disproportionately to the training set (but this point could be due to a misreading on my part of how the table in column 1 and colum 2 is built and how it should be interpreted). Could the authors discuss how this imbalance affects the test results? I would expect under-represented features to show lower performance, and this should be reflected in the evaluation metrics.
The abstract mentions "spinning disk confocal" as an "also included" category. But before that there is no mention of any other modality. It is critical to define all modalities (e.g., widefield vs. confocal) in the introduction/background, as the difference in Z-resolution and PSF (Point Spread Function) means models trained on one may not generalize to the other.
The authors include excitation wavelengths but omit emission wavelengths. SNR and image quality are in some way dependent on the emission filters and camera quantum efficiency at specific wavelengths; therefore, an emission column should be added to the technical tables.
Regarding the layout of Figure 3 I think that for better visual comparison, the figures should be rearranged. Since Restormer is identified as the best-performing model, it should be placed immediately adjacent to the "Ground Truth" (high-SNR) image to allow the reader to easily assess its fidelity.
I'm not sure I got this right but in the abstract the authors mention "12 sub-datasets" but the table contains 15 entries.
I think the sentence " how different stains may impact denoising accuracy" should be rephrased. I can have stains that mark the same structures but with different fluorophores or mechanisms of attachment, and the features will be the same. It is more the target of the stain (which represents the feature content of the image) that would affect how the trained model can be more or less effective on the new target.
In conclusion the subject of the paper is interesting in my opinion and I think that overall the authors did a good job in providing very convincing results and in presenting potential applications of their method. The inclusion of a stitching method for high-megapixel images is also a practical contribution for microscopy applications and quite useful.
-
AbstractBackground Live-cell fluorescence microscopy enables the study of dynamic cellular processes. However, fluorescence microscopy can damage cells and disrupt these dynamic processes through photobleaching and phototoxicity. Reducing light exposure mitigates the effects of photobleaching and phototoxicity but results in low signal-to-noise ratio (SNR) images. Deep learning provides a solution for restoring these low-SNR images. However, these deep learning methods require large, representative datasets for training, testing, and benchmarking, as well as substantial GPU memory, particularly for denoising large images.Results We present a new fluorescence microscopy dataset designed to expand the range of imaging conditions and specimens currently available for evaluating denoising methods. The dataset contains 324 paired …
AbstractBackground Live-cell fluorescence microscopy enables the study of dynamic cellular processes. However, fluorescence microscopy can damage cells and disrupt these dynamic processes through photobleaching and phototoxicity. Reducing light exposure mitigates the effects of photobleaching and phototoxicity but results in low signal-to-noise ratio (SNR) images. Deep learning provides a solution for restoring these low-SNR images. However, these deep learning methods require large, representative datasets for training, testing, and benchmarking, as well as substantial GPU memory, particularly for denoising large images.Results We present a new fluorescence microscopy dataset designed to expand the range of imaging conditions and specimens currently available for evaluating denoising methods. The dataset contains 324 paired high/low-SNR images ranging from four to 282 megapixels across 12 sub-datasets that vary in specimen, objective used, staining type, excitation wavelength, and exposure time. The dataset also includes spinning disk confocal microscopy examples and extreme-noise cases. We evaluated three state-of-the-art deep learning denoising models on the dataset: a supervised transformer-based model, a supervised CNN model, and an unsupervised single image model. We also developed an image stitching method that enables large images to be processed in smaller crops and reconstructed.Conclusions Our dataset provides a diverse benchmark for evaluating deep learning denoising methods, and our stitching method provides a solution to GPU memory constraints encountered when processing large images. Among the evaluated deep learning models, the supervised transformer-based model had the highest denoising performance but required the longest training time.
This work has been peer reviewed in GigaScience(see https://doi.org/10.1093/gigascience/giag071), which carries out single-anonymized peer review. These reviews are published under a CC-BY 4.0 license and were as follows:
Reviewer 1:
This manuscript presents a fluorescence microscopy image dataset intended for evaluating deep learning-based image restoration methods across a range of signal-to-noise ratio (SNR) levels. The study benchmarks three models (Restormer, CARE, and Noise2Fast) and introduces an adaptive image-stitching strategy to accommodate large field-of-view images. Overall, the proposed dataset and the reported benchmark results could be valuable to the community. However, the manuscript would benefit from further clarification and strengthening, particularly regarding the depth and rigor of the experimental analysis. Specific comments are as follows:
The manuscript highlights the novelty and diversity of the proposed dataset, including specimen types that are underrepresented in prior benchmarks as well as examples acquired with spinning-disk confocal microscopy. However, the specific advantages of this dataset over existing resources (e.g., those cited from Zhang et al., Zhou et al., and Hagen et al.) for advancing low-SNR image restoration are not yet sufficiently clear. In particular, the authors should more explicitly articulate what unique challenges the newly included specimen types and the "extreme-noise" cases introduce, and why these cases provide meaningful validation for assessing model robustness and generalization across imaging conditions.
To substantiate a broader conclusion about unsupervised/self-supervised approaches, the authors are encouraged to broaden the benchmark by adding additional representative unsupervised baselines, such as Noise2Void, Noise2Self, Probabilistic Noise2Void, PPNoise2Void, and Self-inspired Noise2Noise, evaluated under the same training and testing protocol.
The performance of the three models is compared under different preprocessing and training pipelines. For a fair and reproducible benchmark, all models should be retrained and evaluated using the same dataset splits and a consistent preprocessing protocol, including data augmentation, normalization schemes, and cropping/tiling strategies. Otherwise, the observed performance differences may reflect implementation choices rather than the intrinsic capabilities of the models. If certain methods impose specific input constraints (e.g., patch size or channel format), the authors should still minimize such discrepancies as much as possible and clearly justify any unavoidable deviations to ensure the comparison is as equitable as possible.
The manuscript refers to "extreme low-SNR" conditions, yet it does not characterize the underlying noise statistics (often well described by a Poisson-Gaussian model in fluorescence microscopy). A clear noise characterization is important for interpreting restoration performance and for selecting or designing appropriate denoising algorithms. The authors should estimate the noise statistics from the acquired measurements (e.g., via variance-mean analysis or other established noise-calibration procedures) or justify why such characterization is not feasible in this study.
-
