?
Optimized Pruning Strategies for Non-Local Architectures in Lung Computed Tomography Tumor Recognition
Attention-augmented encoder-decoder models with non-local modules are effective for lung tumor analysis on computed tomography (CT), yet their computational and memory demands can hinder deployment. We study pruning for a specialized non-local U-Net pipeline that jointly performs neoplasm presence recognition and lesion segmentation from single-channel CT snapshots. We benchmark unstructured and structured pruning baselines and propose a sensitivity-aware, architecture-guided refinement that allocates sparsity across encoder-decoder components and attention modules to better preserve task-critical capacity. On a joint benchmark of public lung CT datasets, the proposed method achieves a superior accuracy-efficiency trade-off. At 50% sparsity, it reduces latency from 41.5 ms to 23.6 ms and FLOPs from 62.4 G to 30.8 G while maintaining strong quality (F1=0.923,mIoU=0.742, Dice=0.847). At 60 % sparsity, it further reduces latency to 19.4ms(FLOPs=23.5G) with stable performance (F1=0.904,mIoU=0.724). Overall, the proposed pruning refinement improves the deployability of attention-based lung CT models without sacrificing clinically meaningful performance.