Sitemap

A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.

Pages

Posts

portfolio

publications

UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling [PDF] [arXiv] [Code] [BibTeX]

Peiming Li, Ziyi Wang, Yulin Yuan, Hong Liu, Xiangming Meng, Junsong Yuan, Mengyuan Liu

In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2025)

A point cloud video is disordered in both space and time, so unfolding it into a 1D sequence hands a state space model neighbours that are not actually related. UST-SSM reorganises the points into semantically coherent sequences by clustering, restores the geometric detail lost in that reordering, and widens the temporal receptive field using non-anchor frames. This carries selective state space models over to point cloud video action recognition and segmentation.

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models [arXiv] [Code] [BibTeX]

Yulin Yuan, Hongshuo Zhao, Xiangming Meng

Accepted to the Conference on Neural Information Processing Systems (NeurIPS 2026)

Parallel decoding in multimodal diffusion large language models can commit several tokens grounded in the same image regions, creating visual redundancy that limits decoding quality. VRCD explicitly measures and controls this redundancy when selecting tokens, encouraging complementary visual grounding and improving multimodal decoding. It yields relative accuracy gains of up to 18.8% on M3CoT and 6.9% on MMBench over confidence-based decoding.

Enhancing Decoupled Posterior Sampling with Data Consistency Guidance for Inverse Problems [Springer] [Code] [BibTeX]

Zhi Qi, Yulin Yuan, Shihong Yuan, Xiangming Meng

In Proceedings of the 22nd International Conference on Intelligent Computing (ICIC 2026), Lecture Notes in Computer Science, pp. 167-179 (Oral Presentation)

Decoupled posterior sampling methods for diffusion inverse problems ignore the measurement during the reverse process, so errors made in the early steps persist and obstruct optimisation later on. GDPS adds a data consistency constraint to that reverse process, which smooths the optimisation trajectory and converges closer to the target distribution. It extends to latent diffusion models and to Tweedie formula, and reaches state-of-the-art results on FFHQ and ImageNet across linear and nonlinear tasks.

One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion [arXiv] [Code] [BibTeX]

Yulin Yuan, Ying Zhang, Xiangming Meng

arXiv preprint arXiv:2609.33698

JPEG-DLM enables efficient diffusion-based text generation by jointly learning compressed embeddings with a flow matching generator and token decoder. Unlike two-stage approaches that freeze a separately trained compression space, joint training adapts the embeddings to both diffusion generation and reliable token decoding while reducing latent sequence length. On LM1B and OWT, JPEG-DLM achieves the lowest mean Gen-PPL and highest throughput among the compared diffusion and flow models.

talks

teaching