Federated Learning for CT-Based Lung Cancer Subtype Classification: A Multi-Dimensional Methodological Comparison

Limited Access
This item is unavailable until:
2028-06-06

Date

2026

Advisors

Journal Title

Journal ISSN

Volume Title

Repository Usage Stats

1
views
0
downloads

Attention Stats

Abstract

\abstract{\textbf{Background:}} Histological subtype classification of lung cancer is clinically important because the major subtypes differ in biological behavior, diagnostic pathways, treatment choice, and prognosis. Computed tomography (CT) is routinely acquired for lung cancer patients to capture tumor-related phenotypic information. CT-based deep learning methods have shown promise in differentiating lung cancer subtypes by data driven learning. Multicenter training offers a favorable route to more accurate and robust models, while direct data sharing across institutions is often limited by privacy concerns and data-governance restrictions. Federated learning (FL) is a privacy-preserving distributed learning mechanism for multicenter data utilization.

\textbf{Purpose:} This study aimed to compare centralized learning (CL) and federated learning for four-class CT-based lung cancer subtype classification, with particular emphasis on examining multiple dimensions of condition, including input image size, prediction strategy, neural network backbone, and data heterogeneity.

\textbf{Materials and Methods:} The dataset was derived from the public Lung-PET-CT-Dx collection released through The Cancer Imaging Archive. After modality filtering and preprocessing, the final analysis set consisted of 820 image-level CT acquisitions, including 606 adenocarcinoma, 80 small cell carcinoma, 10 large cell carcinoma, and 124 squamous cell carcinoma images. Four pretrained deep learning backbones, namely ResNet-50, ConvNeXt-Tiny, Swin-Tiny, and MambaOut-Tiny, were evaluated within a unified five-fold stratified cross-validation framework. Experiments were conducted under two training paradigms, centralized learning and federated learning, two decision strategies, single-model and soft-voting ensemble prediction, two input representations, full-slice and ROI-only. The optimal condition for federated learning under homogenous data conditions was transferred to the examination of heterogeneous data conditions. Heterogeneity in dataset size, image feature, and a mixture of the two were simulated and explored. Federated learning followed a FedAvg-style workflow with three simulated clients. Performance was assessed using macro-F1 as the primary model-selection metric, together with accuracy, balanced accuracy, macro-AUC, sensitivity, specificity, and class-wise metrics.

\textbf{Results:} Under homogenous data condition, the strongest federated result achieved a macro-F1 of $0.8337 \pm 0.0299$ and a macro-AUC of $0.9365 \pm 0.0107$ (single-model, full-slice), was close to the strongest centralized result, which achieved a macro-F1 of $0.8450 \pm 0.0237$ and a macro-AUC of $0.9372 \pm 0.0119$ (single-model, full-slice). Ensemble prediction did not surpass the best individual backbone in macro-F1, but the highest macro-AUC in the study was achieved by the centralized full-slice ACC-weighted ensemble ($0.9687 \pm 0.0134$). Full-slice input consistently outperformed ROI-only input across both centralized and federated settings. Averaged across backbones, full-slice input increased macro-F1 from $0.5435$ to $0.7713$ under centralized learning and from $0.5138$ to $0.7726$ under federated learning. Under several simulated heterogeneous data conditions, federated learning showed in general lower average macro-F1 than centralized learning.

\textbf{Conclusion:} Under homogenous data conditions, FedAvg federated learning maintained competitive classification performance for CT-based lung cancer four-type classification without requiring raw data sharing. Ensemble learning did not improve macro-F1 beyond the best single model, while enhanced probability-based discrimination, particularly macro-AUC. Full-slice CT input consistently outperforms tightly cropped ROI-only input, suggesting that contextual image information may play an important role in multiclass lung cancer subtype classification. Under simulated heterogeneous conditions, FedAvg federated learning showed lower average macro-F1 than centralized learning. Future studies include more comprehensive study of heterogeneous conditions and study of alternative federated learning strategies to reduce the performance gap between FL and CL.

Description

Provenance

Subjects

Medical imaging, deep learning, ensemble learning, federated learning, histological subtype classification, lung cancer

Citation

Citation

Yuan, Yue (2026). Federated Learning for CT-Based Lung Cancer Subtype Classification: A Multi-Dimensional Methodological Comparison. Master's thesis, Duke University. Retrieved from https://hdl.handle.net/10161/35090.

Collections


Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.