Sept. 16, 2022, 1:15 a.m. | Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, Lu Yuan

This paper presents OmniVL, a new foundation model to support both
image-language and video-language tasks using one universal architecture. It
adopts a unified transformer-based visual encoder for both image and video
inputs, and thus can perform joint image-language and video-language
pretraining. We demonstrate, for the first time, such a paradigm benefits both
image and video tasks, as opposed to the conventional one-directional transfer
(e.g., use image-language to help video-language). To this end, we propose a
decoupled joint pretraining of image-language …

