Rethinking Dropout in Transformer-Based Neural Operators: Mechanisms, Effects, and Design Principles

Limited Access
This item is unavailable until:
2027-05-06

Date

2026

Journal Title

Journal ISSN

Volume Title

Attention Stats

Abstract

Transformer models have achieved significant success in natural language processing and computer vision and are now increasingly applied in scientific machine learning to solve complex partial differential equations. However, conventional training techniques, including dropout regularization, are frequently adopted from related fields without a critical evaluation of their appropriateness for partial differential equation (PDE) operator learning. This thesis investigates the effectiveness of dropout in transformer-based operator learning for elliptic PDEs through systematic experiments. Additionally, we introduce a correlated dropout framework to investigate the influence of spatial-frequency characteristics of dropout masks on results, comparing independent, low-frequency, high-frequency, and local Gaussian correlated dropout methods. Our findings consistently show that removing dropout improves accuracy in transformer-based elliptic PDE operator learning. Based on these results, we suggest removing dropout as a practical guideline for designing transformer-based operator learning models.

Description

Provenance

Subjects

Civil engineering

Citation

Citation

Shao, Yuyuan (2026). Rethinking Dropout in Transformer-Based Neural Operators: Mechanisms, Effects, and Design Principles. Master's thesis, Duke University. Retrieved from https://hdl.handle.net/10161/35095.

Collections


Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.