One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion
Jointly learns compressed embeddings with the diffusion generator to enable efficient diffusion-based text generation.
I am a third-year PhD student in the Probabilistic Inference Lab (PIL) at Zhejiang University, expecting to graduate in 2029. I study diffusion models, with a current focus on diffusion language models. I am interested in how parallel denoising can support reliable text generation while reducing the cost of inference.
My recent work focuses on inference-time decisions: which token positions should a model commit together, and how can we guide that choice without retraining? I study how these choices shape parallel generation.
Jointly learns compressed embeddings with the diffusion generator to enable efficient diffusion-based text generation.
Controls visual redundancy during parallel decoding to improve the decoding quality of multimodal diffusion large language models.
Data consistency guidance during the reverse diffusion process for decoupled posterior sampling in inverse problems.
Semantically coherent point sequences and spatio-temporal modeling for point cloud video action recognition and segmentation.
If you are interested in my research or working on related questions, I would be glad to hear from you. I welcome conversations about diffusion models, language generation, and possible collaborations.
You can reach me at yulinyuan@zju.edu.cn.