Lightweight Frequency-Gated Enhancement of nnU-Net Improves Data Efficiency in Nasal Sinus CT Segmentation

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background and Objectives

Automatic segmentation of the nasal cavity and paranasal sinuses from CT aids diagnosis and surgical planning, but clinical datasets in this domain remain small, which can affect training stability and evaluation validity. This study investigates whether the frequency-domain mechanisms of the Adaptive Frequency-Spatial Dual-Stream Network (AFS-DSN) can be transferred, in a lightweight form, into the self-configuring nnU-Net framework, and whether the resulting performance improvement survives rigorous statistical validation. We integrated the AFS-DSN frequency branch and cross-domain attention mechanism into nnU-Net, and corrected a zero-initialization gating deadlock in the original design, using the corrected model (v2) as a common base. On this base, we evaluated three variants: X1, an input-conditioned learnable spectral gate applied to 24 wavelet sub-bands; X2, which adds a boundary-distance training objective; and X3, combining X1 and X2. Experiments used the 130-volume NasalSeg CT dataset, under a full 91/19/20 protocol and a rebuilt, genuine 3-fold cross-validation in which the training set for each fold was reduced by approximately 19%, from 91 to 73–74 cases. Under the full protocol, all variants showed a small but statistically detectable improvement over a matched nnU-Net baseline. X3 achieved a Dice score of 95.78% versus 95.59% for the baseline, an improvement of 0.19 percentage points (95% CI [0.09, 0.29]; Holm-adjusted p = 0.004), and X3 also reduced average surface distance (ASD) from 0.252 mm to 0.235 mm. X3 was statistically equivalent to either individual component within a margin of ± 0.10 percentage points, suggesting a shared performance ceiling rather than a complementary gain. Under reduced training data, baseline Dice dropped by 2.71 percentage points (95.59% to 92.88%), while X3 remained essentially unchanged (95.78% to 95.75%), a cross-fold advantage of +2.88 percentage points (fold-level 95% CI [0.80, 4.96]) that held across all 20 test cases (sign test, p = 1.9 × 10 −6 ) and was corroborated by average surface distance. The practical value of this combined approach lies primarily in within-distribution data efficiency, rather than in maximizing peak segmentation accuracy.

Article activity feed