UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling [PDF] [arXiv] [Code] [BibTeX]
In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2025)
A point cloud video is disordered in both space and time, so unfolding it into a 1D sequence hands a state space model neighbours that are not actually related. UST-SSM reorganises the points into semantically coherent sequences by clustering, restores the geometric detail lost in that reordering, and widens the temporal receptive field using non-anchor frames. This carries selective state space models over to point cloud video action recognition and segmentation.
