Advancing Multimodal Generative Models via Structured Representation Learning and Semantic Planning

Limited Access
This item is unavailable until:
2027-05-06

Date

2026

Journal Title

Journal ISSN

Volume Title

Attention Stats

Abstract

Modern multimodal generative models produce visually stunning images and videos, yet they frequently fail to follow complex instructions, maintain semantic coherence over long temporal horizons, or satisfy user-specified constraints such as privacy. These shortcomings stem from a common architectural bottleneck: monolithic end-to-end generators conflate high-level semantic understanding with low-level pixel synthesis, offering no explicit mechanism for abstraction, hierarchical reasoning, or constraint enforcement.

This dissertation demonstrates that structured intermediate representations and explicit semantic planning can close this gap. We pursue two complementary thrusts---learning representations whose internal organization reflects semantic structure, and delegating high-level reasoning to dedicated planning modules---and develop them across five interconnected works that trace a progression from foundational tokenization to planning-guided, constraint-aware generation.

We begin by showing that discrete video tokens, generated autoregressively with a three-dimensional sparse attention mechanism and pretrained on over 136 million text--video pairs, enable one of the first open-domain text-to-video models (GODIVA). We then reveal that the ordering within such discrete codes matters: Progressive Quantization VAE (PQ-VAE) learns a hierarchy of tokens ranked by semantic importance, where the leading codes capture global structure and subsequent codes add fine-grained detail, yielding superior compression and interpretability. Extending the principle of structured representation to continuous embeddings, we adapt Orthogonal Low-rank Embedding for self-supervised learning (SSOLE), overcoming collapse and sign-ambiguity barriers to achieve competitive performance on ImageNet without large batches, memory banks, or dual encoders.

Building on these representation foundations, we introduce Plan-X, a framework that decouples semantic planning from visual synthesis: a multimodal language model produces structured spatio-temporal semantic tokens that guide a video Diffusion Transformer, dramatically reducing visual hallucination and improving instruction alignment. Finally, we demonstrate that the same representation-learning principles enable constraint-driven generation for privacy: an ID-augmented contrastive encoder, trained with the SSOLE objective on synthetic face-swap variants, learns identity-agnostic attribute codes that condition a diffusion generator to produce anonymized faces preserving expression, gaze, and pose while effectively erasing identity.

Extensive experiments across video generation, image reconstruction, self-supervised learning, instruction-following synthesis, and privacy-preserving face anonymization confirm the central thesis: structured intermediate representations and explicit planning are essential for advancing multimodal generative models from visually plausible output toward faithful, controllable, and responsible generation.

Description

Provenance

Subjects

Computer engineering

Citation

Citation

Huang, Lun (2026). Advancing Multimodal Generative Models via Structured Representation Learning and Semantic Planning. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35189.

Collections


Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.