Intensity-Stratified Evaluation of Brain CT-to-T1 MRI Synthesis Using Deep Learning Models
Date
2026
Authors
Advisors
Journal Title
Journal ISSN
Volume Title
Attention Stats
Abstract
Consider two CT-to-T1 MRI synthesis models evaluated on the same 425 validation cases: one produces spatially diffuse, contrast-attenuated outputs; the other exhibits pronounced fluctuation artifacts, with markedly elevated error at air, CSF, and bone boundaries, as well as white matter. By global PSNR, they appear indistinguishable—18.58 versus 18.57 dB. This numerical coincidence is not merely a curiosity; it suggests a potential limitation in how medical image synthesis is evaluated.CT to T1 MRI synthesis has become increasingly relevant for radiation therapy planning, multimodal registration, and clinical scenarios where MRI acquisition is unavailable or impractical. While recent deep learning models have shown promising results, evaluation practice has remained largely unchanged: most studies still rely on global image-level metrics such as peak signal-to-noise ratio, structural similarity, and mean absolute error. These metrics aggregate errors across the entire image and can easily miss failures concentrated in specific tissue tiers—failures that may be clinically significant precisely because they are localized. This thesis addresses that gap by developing a region-aware, HU-based intensity-stratified evaluation framework for CT to T1 MRI synthesis. We combine global and brain-masked quantitative metrics with Hounsfield-unit-based tissue stratification and HU-annotated line profile analysis. Five representative models are evaluated under matched preprocessing and evaluation conditions: VAE, U-Net, Pix2Pix, CycleGAN, and DiffMa. Supervised paired models—Pix2Pix and U-Net in particular—achieve the most consistent fidelity across tissue types. A key finding is a previously unreported SSIM ranking reversal: CycleGAN ranks above VAE on global SSIM (0.675 vs. 0.662), but the ranking inverts when evaluation is restricted to the brain mask (VAE 0.690 vs. CycleGAN 0.685). This reversal exposes a potential limitation of SSIM—the metric can yield different model rankings depending on whether evaluation is performed globally or within the brain mask, because spatial smoothness and background inclusion may be weighted differently from within-brain intensity fidelity. The framework also linked these failures to plausible architectural mechanisms. VAE's behavior is consistent with KL-bottleneck-driven suppression of high-frequency intensity variation, which can inflate SSIM even as tissue contrast is attenuated. CycleGAN's behavior is consistent with instability at high-contrast boundaries under the cycle-consistency constraint, manifesting as fluctuation artifacts that paired supervision did not fully suppress in my experiments. DiffMa showed high variability across cases; given the stochastic nature of DDPM sampling, repeated-run variability is a plausible concern, though it was not directly quantified in this study. In the bone tier, Pix2Pix achieved lower MAE than DiffMa in all or nearly all of the 357 evaluated cases (Wilcoxon W=0, p<0.001)—a near-unanimous advantage that highlights the extent of diffusion-based synthesis errors at high-density tissue boundaries. Tissue-specific MAE computed within HU-defined tissue tiers reveals that models with nearly identical global scores can fail in qualitatively different ways across CSF, white matter, gray matter, and bone regions—differences with direct consequences for radiation therapy planning, automated tissue segmentation, and longitudinal volumetric measurement.
Type
Description
Provenance
Subjects
Citation
Permalink
Citation
Zhou, Wen (2026). Intensity-Stratified Evaluation of Brain CT-to-T1 MRI Synthesis Using Deep Learning Models. Master's thesis, Duke University. Retrieved from https://hdl.handle.net/10161/35063.
Collections
Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.
