Web: http://arxiv.org/abs/2205.03569

June 17, 2022, 1:13 a.m. | Bing Li, Jiaxin Chen, Dongming Zhang, Xiuguo Bao, Di Huang

cs.CV updates on arXiv.org arxiv.org

Compressed video action recognition has recently drawn growing attention,
since it remarkably reduces the storage and computational cost via replacing
raw videos by sparsely sampled RGB frames and compressed motion cues (e.g.,
motion vectors and residuals). However, this task severely suffers from the
coarse and noisy dynamics and the insufficient fusion of the heterogeneous RGB
and motion modalities. To address the two issues above, this paper proposes a
novel framework, namely Attentive Cross-modal Interaction Network with Motion
Enhancement (MEACI-Net). It …

arxiv cross cv learning representation representation learning video

