Calibrated LRT Guidance for Offline Diffusion Policies

dc.contributor.advisor

Cheng, Xiang XC

dc.contributor.author

Sun, Ximan

dc.date.accessioned

2026-07-06T19:50:07Z

dc.date.available

2026-07-06T19:50:07Z

dc.date.issued

2026

dc.department

Electrical and Computer Engineering

dc.description.abstract

Diffusion policies are competitive for offline Reinforcement Learning but are typically guided at sampling time by heuristics that lack a statistical notion of risk. We introduce LRT-Diffusion, a risk-aware sampling rule that performs evidence accumulation between two inference-time heads: an unconditional background head and a state-conditional good head. Concretely, we accumulate a log-likelihood ratio and gate the conditional mean with a logistic controller whose threshold $\tau$ is calibrated once per task and per sampler under $H_0$ to meet a user-specified Type-I level $\alpha$. This turns guidance from a fixed push into an \emph{evidence-driven} adjustment with a user-interpretable risk budget. Importantly, we deliberately leave training vanilla (two heads with standard $\epsilon$-prediction) under the structure of DDPM. LRT guidance composes naturally with Q-gradients: critic-gradient updates can be taken at the unconditional mean, at the LRT-gated mean, or a blend, exposing a continuum from exploitation to conservatism. We standardize states/actions consistently at train and test time and report a state-conditional OOD metric alongside return. On D4RL MuJoCo tasks, LRT-Diffusion yields a calibrated return–risk frontier: LRT often reduces state-conditional OOD, and combining with a small Q-step increases return along the frontier. Theoretically, we establish level-$\alpha$ calibration, stability bounds, and a return comparison showing when evidence-gated guidance is preferable to pure Q-guidance. Overall, LRT-Diffusion is a drop-in, inference-time method that adds principled, calibrated risk control to diffusion policies for offline RL.

dc.identifier.uri

https://hdl.handle.net/10161/35064

dc.rights.uri

https://creativecommons.org/licenses/by-nc-nd/4.0/

dc.subject

Computer engineering

dc.subject

Computer science

dc.subject

Diffusion Policies

dc.subject

Inference

dc.subject

Likelihood‑ratio Test

dc.subject

Offline Reinforcement Learning

dc.subject

Out‑of‑distribution Detection

dc.subject

Risk‑aware Guidance

dc.title

Calibrated LRT Guidance for Offline Diffusion Policies

dc.type

Master's thesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Sun_duke_0066N_19334.pdf
Size:
1.19 MB
Format:
Adobe Portable Document Format

Collections