Efficient Learning from Weak Supervision: From Data and Model Reuse to Agentic Autonomy
Date
2026
Authors
Advisors
Journal Title
Journal ISSN
Volume Title
Repository Usage Stats
views
downloads
Attention Stats
Abstract
As machine learning systems transition from narrow, task-specific models to general-purpose agents, the primary bottleneck remains the high cost of supervision. This dissertation investigates the paradigm of \emph{efficiently} solving \emph{new} learning problems under increasingly \emph{weak} supervision, tracing a technical evolution from data and model reuse to the development of autonomous agentic systems.
The first part of this work addresses \emph{learning from shifted data}, focusing on decision-making where the available data does not match the target environment. I present frameworks for (i) \emph{transfer learning in causal inference tasks}, which reuse data, model, and knowledge across disparate domains, and (ii) \emph{offline assortment optimization}, which must account for the distributional shift between historical behavior policies and the optimal target policy. Through delicate algorithmic designs, we demonstrate that structured data reuse can solve complex decision problems without the need for expensive, real-world exploration.
Building on these foundations, the second part explores robustness as a guaranteed generalization challenge. I present two distinct approaches: (i) a \emph{data augmentation} framework for causal inference and (ii) a \emph{robust reinforcement learning} (RL) algorithm. By framing robustness as the ability to maintain performance across a set of unknown target environments, these methods provide bridges to supervision even \emph{weaker} than one considered in the first part. Specifically, robustness requires the models to generalize (perform well) without information about the specific deployment environment at training time.
The third part shifts toward In-Context Reinforcement Learning (ICRL), where a single model is trained to solve entirely unseen RL tasks through its context window. I introduce three key contributions in this space. First, I present a framework for \emph{ICRL using suboptimal historical trajectories}. Through this framework, we circumvent the requirement to collect high-quality expert data for pretraining ICRL models. Second, I introduce a paradigm for \emph{ICRL using only preference-based feedback}, bypassing the need for explicit reward signals. In particular, under this novel paradigm, both pretraining and deployment of ICRL models require only a few preference labels, significantly enhancing the scalability of ICRL applications. We investigate both step and trajectory preference feedback and propose distinct preference-native ICRL algorithms for difference feedback. Our technical highlights here include a direct policy optimization algorithm using step preference and an in-context state-action creditor pretrained from only trajectory preference. Finally, I investigate the \emph{theoretical connection between generalization and robustness in in-context learners}: I demonstrate how the inherent generalization capabilities of large-scale models can be leveraged to achieve and efficiently enhance robustness in unseen tasks.
Lastly, I present the development of an \emph{agentic system for GPU kernel optimization}. This work integrates LLM-based reasoning with low-level hardware constraints, demonstrating how agentic architectures can bridge the gap between high-level intent and complex hardware execution.
Collectively, the contributions of this dissertation provide a roadmap for building AI systems that are data-efficient, robust by design, and capable of autonomous adaptation under sparse supervision.
Type
Department
Description
Provenance
Subjects
Citation
Permalink
Citation
Dong, Juncheng (2026). Efficient Learning from Weak Supervision: From Data and Model Reuse to Agentic Autonomy. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35317.
Collections
Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.
