UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling
In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2025)
Selective state space models handle long sequences at linear cost, which makes them attractive for video. Point cloud video resists that treatment: the points carry no order in space and no consistent identity across time, so unfolding a clip into a 1D sequence by scanning frames in order gives the model a sequence whose neighbouring elements are not actually related.
UST-SSM addresses this in three parts. Spatio-temporal selection scanning reorganises the unordered points into semantically coherent sequences through clustering, so that adjacent positions in the sequence are adjacent in meaning. Spatio-temporal structure aggregation recovers the geometric detail lost during that reordering. Temporal interaction sampling widens the receptive field by drawing on non-anchor frames, which strengthens the dependencies across time.
The method is evaluated on point cloud video benchmarks for action recognition and segmentation.
