Calibrated LRT Guidance for Offline Diffusion Policies

Loading...

Date

2026

Journal Title

Journal ISSN

Volume Title

Attention Stats

Abstract

Diffusion policies are competitive for offline Reinforcement Learning but are typically guided at sampling time by heuristics that lack a statistical notion of risk. We introduce LRT-Diffusion, a risk-aware sampling rule that performs evidence accumulation between two inference-time heads: an unconditional background head and a state-conditional good head. Concretely, we accumulate a log-likelihood ratio and gate the conditional mean with a logistic controller whose threshold $\tau$ is calibrated once per task and per sampler under $H_0$ to meet a user-specified Type-I level $\alpha$. This turns guidance from a fixed push into an \emph{evidence-driven} adjustment with a user-interpretable risk budget. Importantly, we deliberately leave training vanilla (two heads with standard $\epsilon$-prediction) under the structure of DDPM. LRT guidance composes naturally with Q-gradients: critic-gradient updates can be taken at the unconditional mean, at the LRT-gated mean, or a blend, exposing a continuum from exploitation to conservatism. We standardize states/actions consistently at train and test time and report a state-conditional OOD metric alongside return. On D4RL MuJoCo tasks, LRT-Diffusion yields a calibrated return–risk frontier: LRT often reduces state-conditional OOD, and combining with a small Q-step increases return along the frontier. Theoretically, we establish level-$\alpha$ calibration, stability bounds, and a return comparison showing when evidence-gated guidance is preferable to pure Q-guidance. Overall, LRT-Diffusion is a drop-in, inference-time method that adds principled, calibrated risk control to diffusion policies for offline RL.

Description

Provenance

Subjects

Computer engineering, Computer science, Diffusion Policies, Inference, Likelihood‑ratio Test, Offline Reinforcement Learning, Out‑of‑distribution Detection, Risk‑aware Guidance

Citation

Citation

Sun, Ximan (2026). Calibrated LRT Guidance for Offline Diffusion Policies. Master's thesis, Duke University. Retrieved from https://hdl.handle.net/10161/35064.

Collections


Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.