One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion

Yulin Yuan, Ying Zhang, Xiangming Meng

arXiv preprint arXiv:2609.33698

JPEG-DLM enables efficient diffusion-based text generation by jointly learning compressed embeddings with the generative model. Compressing multiple tokens into fewer latent positions reduces the sequence length processed at every diffusion step and therefore lowers generation cost.

Existing two-stage methods train the compression space first and then freeze it while training the diffusion model. JPEG-DLM instead jointly trains the compressor, flow matching model, and decoding module. Joint-embedding prediction shapes compressed embeddings that are well suited to diffusion generation and can still be decoded reliably into tokens.

Across LM1B and OWT, JPEG-DLM achieves the lowest mean Gen-PPL and highest throughput among the compared diffusion and flow models. At a compression rate of 0.5 on OWT, it reaches a Gen-PPL of 34.52 at approximately 2.3x the throughput of ELF.